Senior Software Engineer, Quantized Inference

NVIDIA NVIDIA · Semiconductors · Redmond, WA +1

Senior Software Engineer focused on optimizing LLM inference through quantization and sparsity techniques, implementing recipes in inference engines like vLLM and TRT-LLM, and developing model export pipelines.

What you'd actually do

  1. Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang)
  2. Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream serving
  3. Build prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimization
  4. Develop data analysis tooling and visualizations for numerics debugging
  5. Improve developer productivity across the team: CI, build systems, training infrastructure, pipeline friction

Skills

Required

  • Python
  • C++
  • Software engineering fundamentals
  • ML accelerators
  • PyTorch internals
  • large open-source codebase
  • written and verbal communication

Nice to have

  • Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang)
  • Triton kernel development
  • Debugging numerical issues across mixed-precision boundaries
  • Model compression techniques: PTQ, QAT, structured/unstructured sparsity

What the JD emphasized

  • Quantized Inference
  • LLMs
  • inference engines
  • vLLM
  • TRT-LLM
  • SGLang
  • ModelOpt
  • Megatron-LM
  • quantize/dequantize
  • MoE layers
  • throughput
  • interactivity
  • productization effort
  • Python
  • C++
  • ML accelerators
  • PyTorch internals
  • large open-source codebase
  • move fast with ambiguous requirements

Other signals

  • LLM inference optimization
  • quantization
  • kernel development
  • model serving