Senior Software Engineer - GPU Local AI Platforms

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +4

Senior Software Engineer focused on optimizing LLM inference performance on NVIDIA edge AI hardware. The role involves tracking innovations in inference frameworks, analyzing model architectures and algorithms, characterizing multi-node inference, producing performance reports, and owning the model validation and recipe development workflow for new model releases.

What you'd actually do

  1. Track and evaluate innovations in leading open-source LLM inference frameworks — identify performance-critical features and algorithmic improvements relevant to NVIDIA edge AI hardware
  2. Analyze how new model architectures and inference algorithms (attention variants, MoE routing, speculative decoding, multi-token prediction, quantized inference) map onto NVIDIA GPU architecture — identify mismatch, fallback paths, and optimization opportunities
  3. Characterize multi-node inference behavior: collective communication primitives (NCCL/RCCL), topology-aware all-reduce strategies, and parallelism efficiency on edge cluster configurations
  4. Produce performance analysis reports mapping theoretical hardware limits (memory bandwidth, FLOP/s, interconnect throughput) to observed inference throughput, latency, and utilization
  5. Own the model validation workflow for new model releases: architecture compatibility assessment, inference recipe development, performance characterization, and publication to developer recipe sites

Skills

Required

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • 12+ years of software engineering with depth in GPU computing, ML systems, or high-performance inference
  • Strong Python or C++ programming, software design, and software engineering skills.
  • Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) — you understand how thread blocks, memory hierarchy, and warp execution affect real-world performance
  • Working knowledge of LLM inference internals: attention mechanisms, KV-cache management, continuous batching, quantization formats, and tensor parallelism
  • Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization, runtime configuration, NVIDIA Container Toolkit
  • Strong analytical skills: ability to form a performance hypothesis, design an experiment, interpret results, and communicate findings clearly

Nice to have

  • Track and evaluate innovations in leading open-source LLM inference frameworks
  • Analyze how new model architectures and inference algorithms map onto NVIDIA GPU architecture
  • Characterize multi-node inference behavior
  • Produce performance analysis reports
  • Own the model validation workflow for new model releases
  • Develop and maintain developer-facing inference recipes
  • Engage with community and partners on model bring-up questions

What the JD emphasized

  • GPU kernel development or optimization (CUDA/C++, Triton, or equivalent)
  • LLM inference internals
  • performance hypothesis
  • experiment
  • interpret results
  • communicate findings clearly

Other signals

  • LLM inference frameworks
  • NVIDIA edge AI hardware
  • performance optimization
  • model bring-up infrastructure
  • developer-facing inference recipes