Senior System Software Engineer - Localai

NVIDIA NVIDIA · Semiconductors · Pune, India

Senior Systems Software Engineer to develop efficient on-device AI software for RTX and DGX-class systems, focusing on high-performance local inference with low latency and optimized memory utilization.

What you'd actually do

  1. Build and optimize the local AI inference stack for RTX, RTX Pro, and DGX GPUs, with a focus on performance, stability, and scalability across diverse hardware architectures.
  2. Design and develop modern inference runtimes and execution stacks using frameworks such as llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, supporting LLM, vision-language, TTS, ASR, and diffusion-based AI workloads.
  3. Perform end-to-end optimization of AI models, data pipelines, and inference runtimes to maximize performance on current and next-generation GPU architectures.
  4. Apply model optimization techniques, including quantization, pruning, sparsity, and distillation, to enable efficient deployment of large models on local and edge devices.
  5. Conduct system-level debugging, performance tuning, and performance-accuracy trade-off analysis; develop infrastructure for performance and accuracy sweeps; analyse results to identify gaps and drive fixes; and establish engineering guidelines to accelerate bring-up and ensure production readiness of new models and inference backends.

Skills

Required

  • C++
  • Data Structures
  • Algorithms
  • Machine Learning
  • AI inference pipelines
  • Llama.cpp
  • vLLM
  • PyTorch
  • Windows ML
  • DXCGC
  • TensorRT
  • Inference backends
  • Runtime internals
  • Scheduling
  • Memory management
  • KV-cache behaviour
  • Graph execution
  • Quantization
  • Hardware-aware optimization
  • System-level debugging
  • Performance tuning
  • Performance-accuracy trade-off analysis
  • Communication skills

Nice to have

  • Modern machine learning
  • Deep neural network
  • Generative AI techniques
  • Open-source contributions
  • Low-level system programming
  • GPU programming
  • CUDA
  • High-performance systems development
  • Open-source inference runtimes
  • Model tooling
  • Performance infrastructure

What the JD emphasized

  • excellent C++ programming and debugging skills
  • proven experience developing and optimizing AI inference pipelines and applications
  • deep understanding of inference backends and runtime internals
  • strong analytical and problem-solving skills

Other signals

  • Optimizing inference stack
  • Low-latency local inference
  • On-device AI software