Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +1 · Remote

NVIDIA is seeking a Senior Inference Performance Engineer to optimize AI inference benchmarks using an autonomous optimization framework powered by AI agents. The role involves improving throughput and interactivity of AI workloads by exploring various optimization techniques and serving architectures, profiling performance, and collaborating with different teams to implement improvements.

What you'd actually do

  1. Distill your performance instincts into reusable skills, workflows, and evidence-backed methodologies that AI agents can complete autonomously. Review agent-generated experiments, validate findings, and curate best-known configurations.
  2. Performance improvement of AI inference workloads that methodically increase throughput-per-GPU and user interactivity by exploring configuration options, parallelism techniques, batching, KV cache handling, quantization, and speculative decoding settings.
  3. Measure and optimize both aggregated and disaggregated serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA's latest GPU platforms.
  4. Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find credible headroom and drive fixes from hypothesis to measured wins.
  5. Land improvements upstream: serving framework patches, optimized kernels, and deployment recipes that advance the public Pareto frontier while maintaining strict model correctness.

Skills

Required

  • AI model execution optimization
  • continuous batching
  • throughput-latency tradeoffs
  • KV cache and memory limitations
  • parallel processing techniques
  • MoE serving
  • quantization
  • serving SLAs
  • benchmarking GPU workloads
  • profiling GPU workloads
  • Nsight Systems
  • Nsight Compute
  • CUPTI
  • PyTorch profiler
  • kernel-level performance data interpretation
  • Python engineering
  • C++
  • CUDA
  • serving codebases
  • experimental methodology
  • controlled single-variable comparisons
  • reproducible benchmarks
  • evidence-backed optimization decisions
  • written and verbal communication

Nice to have

  • TensorRT-LLM contributions
  • vLLM contributions
  • SGLang contributions
  • FlashInfer contributions
  • Dynamo contributions
  • disaggregated serving
  • wide expert-parallel MoE inference
  • KV cache transfer
  • NCCL
  • NIXL
  • NVSHMEM
  • multi-node scale
  • CUDA kernel authorship
  • optimization experience on Hopper/Blackwell architectures
  • Tensor Cores
  • TMA
  • warp specialization
  • MLPerf Inference
  • SemiAnalysis InferenceX
  • building agentic AI workflows
  • operating agentic AI workflows

What the JD emphasized

  • extensive knowledge of the efficiency and optimization involved in AI model execution
  • hands-on experience benchmarking and profiling GPU workloads
  • Strong Python engineering skills
  • ability to navigate and modify large C++/CUDA serving codebases
  • Rigorous experimental methodology
  • controlled single-variable comparisons
  • reproducible benchmarks
  • evidence-backed optimization decisions
  • autonomous optimization framework
  • AI agents use this framework

Other signals

  • optimize AI inference
  • autonomous optimization framework
  • AI agents use this framework
  • performance optimization