Senior System Software Engineer – Dynamo Tools

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +4 · Remote

Senior System Software Engineer to lead the design, development, and roadmap of AI-Perf, NVIDIA's framework for benchmarking, experimentation, and analysis of LLMs, Generative AI, and deep learning inference workloads. This role involves building scalable features for performance measurement and integrating with various inference stacks.

What you'd actually do

  1. Lead the design, development, and roadmap of AI-Perf, defining benchmarking methodologies, performance metrics, and reproducible experimental workflows.
  2. Build scalable and high-performance features to measure latency, throughput, and efficiency across AI models and distributed systems.
  3. Partner with AI researchers, platform teams, and engineers to translate experimental challenges into robust, user-friendly performance tooling.
  4. Integrate AI-Perf with the Dynamo Inference Stack, other NVIDIA inference stacks, and open-source inference frameworks, delivering end-to-end performance insights for researchers and production users.

Skills

Required

  • Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, or related field—or equivalent experience.
  • 3+ years of experience in systems software, distributed performance engineering, or AI infrastructure research.
  • Expert-level Python skills, including profiling, optimization, automation, and debugging of complex systems.
  • Deep knowledge of distributed systems concepts, including scalability, concurrency, fault tolerance, and performance trade-offs.

Nice to have

  • Experience designing or maintaining performance benchmarking frameworks or tooling for AI/ML systems.
  • Hands-on experience with LLMs and deep learning frameworks such as PyTorch, TensorFlow, TensorRT, or ONNX Runtime.
  • Contributions to open-source or research projects in AI performance, infrastructure, or distributed systems.
  • Experience running large-scale inference experiments across cloud and on-prem environments (AWS, Azure, GCP, bare metal).

What the JD emphasized

  • Lead the design, development, and roadmap
  • AI-Perf
  • benchmarking methodologies
  • performance metrics
  • experimental workflows
  • scalable and high-performance features
  • measure latency, throughput, and efficiency
  • AI models and distributed systems
  • Partner with AI researchers, platform teams, and engineers
  • translate experimental challenges into robust, user-friendly performance tooling
  • Integrate AI-Perf with the Dynamo Inference Stack, other NVIDIA inference stacks, and open-source inference frameworks
  • end-to-end performance insights
  • researchers and production users
  • 3+ years of experience in systems software, distributed performance engineering, or AI infrastructure research
  • Expert-level Python skills, including profiling, optimization, automation, and debugging of complex systems
  • Deep knowledge of distributed systems concepts, including scalability, concurrency, fault tolerance, and performance trade-offs
  • Experience designing or maintaining performance benchmarking frameworks or tooling for AI/ML systems
  • Hands-on experience with LLMs and deep learning frameworks such as PyTorch, TensorFlow, TensorRT, or ONNX Runtime
  • Contributions to open-source or research projects in AI performance, infrastructure, or distributed systems
  • Experience running large-scale inference experiments across cloud and on-prem environments (AWS, Azure, GCP, bare metal)

Other signals

  • AI-Perf framework for LLMs and Generative AI
  • benchmarking, experimentation, and analysis of deep learning inference workloads
  • scalable and high-performance features to measure latency, throughput, and efficiency across AI models and distributed systems
  • Integrate AI-Perf with the Dynamo Inference Stack, other NVIDIA inference stacks, and open-source inference frameworks