Software Engineer, Cuda Deep Learning Systems

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +2 · Remote

Software Engineer focused on optimizing deep learning systems at the intersection of CUDA and high-level DL frameworks. The role involves research, prototyping, and optimizing distributed computing systems, custom CUDA kernels, and analyzing hardware-software interactions for both training and inference pipelines. Emphasis on squeezing performance from accelerator architectures for emerging AI workloads.

What you'd actually do

  1. Explore, research, and prototype novel systems optimizations for advanced deep learning models at the intersection of high-level DL frameworks and low-level CUDA through modeling, simulation, and silicon prototyping.
  2. Architect and optimize distributed computing systems that scale seamlessly from a single node to massive, cluster-scale supercomputing environments.
  3. Design, implement, and optimize custom high-performance CUDA kernels tailored to emerging neural network architectures and workloads.
  4. Analyze complex hardware-software interactions to identify and resolve performance bottlenecks in both training and inference pipelines.
  5. Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts to co-design systems and algorithms that improve accelerator compute utilization, memory bandwidth, cross-node network communication efficiency and programmability.

Skills

Required

  • C++
  • Python
  • Deep Learning fundamentals
  • transformers
  • distributed computing principles
  • multi-node scaling
  • systems programming
  • computer architecture
  • low-level systems performance optimization
  • CUDA programming
  • kernel optimization
  • workload profiling
  • generative AI models
  • large language models
  • machine learning systems

Nice to have

  • PyTorch
  • JAX
  • TensorRT
  • vLLM
  • sgLang
  • Nemo
  • Megatron
  • NCCL
  • MPI
  • UCX
  • pipeline parallelism
  • tensor parallelism
  • expert parallelism
  • NVFP4
  • MXFP4
  • FP8
  • INT8
  • Triton
  • XLA
  • torch.compile
  • agentic AI systems

What the JD emphasized

  • low-level hardware optimization has never been more critical
  • unlock maximum hardware performance
  • highly technical group exploring uncharted territories
  • squeezing every ounce of performance
  • low-level CUDA
  • custom high-performance CUDA kernels
  • performance bottlenecks
  • accelerator compute utilization
  • low-level systems performance optimization
  • CUDA programming, kernel optimization, and workload profiling
  • performance internals
  • low-precision arithmetic

Other signals

  • optimization
  • performance
  • distributed systems
  • CUDA
  • deep learning frameworks