Manager, System Software Engineering - Local AI

NVIDIA NVIDIA · Semiconductors · Pune, India

Manager for an on-device AI inference platform team, focusing on high-performance local inference, agentic workloads, low latency, efficient memory use, and scalable infrastructure for RTX, RTX Pro, and DGX GPUs. The role involves leading a team, driving cross-functional alignment, providing technical leadership for inference runtimes, mentoring engineers, and coordinating end-to-end optimization of AI models and inference pipelines.

What you'd actually do

  1. Lead and grow a team building the on-device AI inference platform for RTX, RTX Pro, and DGX GPUs, with accountability for execution, technical direction, delivery quality, and roadmap alignment.
  2. Drive cross-functional alignment with NVIDIA’s software, research, architecture, and product teams, along with industry partners and open-source communities, to build strategy and strengthen the AI ecosystem across RTX and DGX platforms.
  3. Provide technical leadership for the architecture and evolution of modern inference runtimes and execution stacks across frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, spanning workloads including LLMs, vision-language models, TTS, ASR, and diffusion models.
  4. Mentor engineers, develop technical leaders, and foster a high-performance team culture centred on innovation, collaboration, and operational excellence.
  5. Coordinate end-to-end optimization of AI models, data pipelines, and inference runtimes to improve performance across current and next-generation GPU architectures.

Skills

Required

  • Leadership experience in systems software, AI infrastructure, or inference runtimes
  • C++ software development
  • Debugging
  • Data structures
  • Algorithms
  • Machine learning systems
  • AI inference pipelines
  • Deep Learning frameworks (Llama.cpp, vLLM, PyTorch, WinML, DXCGC, TensorRT)
  • Inference backends and runtime internals (scheduling, memory management, KV-cache, graph execution, quantization, hardware-aware optimization)
  • Analytical and problem-solving skills
  • Communication skills

Nice to have

  • Modern machine learning
  • Deep neural networks
  • Generative AI
  • Contributions to notable open-source projects
  • Building teams
  • Setting technical vision
  • Scaling execution
  • Lower-level systems or GPU programming (CUDA)
  • High-performance systems development
  • Contributions to open-source inference runtimes, model tooling, or performance infrastructure
  • Experience with frameworks and APIs including Llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, and vLLM

What the JD emphasized

  • 5+ overall years of industry experience and 2+ years of engineering leadership experience
  • Strong technical foundation in C++ software development, debugging, data structures, algorithms, and machine learning systems.
  • Extensive background in AI inference pipelines and Deep Learning frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
  • Deep understanding of inference backends and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.

Other signals

  • on-device AI inference platform
  • local execution
  • low latency
  • efficient memory use
  • scalable infrastructure
  • resource-constrained platforms