Senior Inference Engineer, GPU Kernel Optimization

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +4

Senior Inference Engineer focused on GPU kernel optimization for LLM inference, developing benchmarking infrastructure, performance projection tooling, and agentic optimization systems to improve performance at the assembly layer. Collaborates with compiler, kernel, hardware, and framework teams to deliver measurable gains.

What you'd actually do

  1. The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack.
  2. The first is GPU kernel microbenchmarking: measuring competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand.
  3. The second is end-to-end model performance analysis: connecting performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments.
  4. The third is agentic kernel optimization: applying AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements.
  5. All three streams converge in close collaboration with compiler, hardware, kernel, and framework teams to deliver upstream improvements and production-grade performance gains.

Skills

Required

  • Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.
  • Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.

Nice to have

  • Deep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).
  • Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.
  • Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).

What the JD emphasized

  • agentic optimization systems
  • agentic kernel optimization
  • agentic AI systems
  • agentic systems end-to-end

Other signals

  • LLM inference performance optimization
  • GPU kernel optimization
  • agentic optimization systems