Engineering Manager, Deep Learning Inference

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +5 · Remote

Engineering Manager to lead a team focused on deep learning inference and GPU-accelerated software, advancing open-source frameworks like SGLang and vLLM for scalable and efficient AI deployment on NVIDIA GPUs.

What you'd actually do

  1. Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
  2. Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering.
  3. Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.
  4. Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.
  5. Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).

Skills

Required

  • C/C++ software design and development
  • Python
  • GPU programming (CUDA, Triton, CUTLASS)
  • performance optimization
  • deploying or optimizing deep learning models in production environments
  • leading teams using Agile or collaborative software development practices

Nice to have

  • Significant open-source contributions to deep learning or inference frameworks such as PyTorch, vLLM, SGLang, Triton, or TensorRT-LLM
  • Deep understanding of multi-GPU communications (NIXL, NCCL, NVSHMEM) and distributed inference architectures
  • Expertise in performance modeling, profiling, and system-level optimization across CPU and GPU platforms
  • Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects with measurable impact
  • Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering

What the JD emphasized

  • deep learning inference
  • GPU-accelerated software
  • open-source frameworks
  • LLM
  • multimodal
  • generative AI
  • NVIDIA GPUs
  • performance tuning
  • profiling
  • optimization
  • CUDA
  • Triton
  • CUTLASS
  • multi-GPU communications

Other signals

  • leading a team
  • deep learning inference
  • GPU-accelerated software
  • open-source frameworks
  • LLM, multimodal, and generative AI applications