Senior AI Training Performance Architect

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA

Senior engineer focused on performance analysis and optimization of AI training workloads across the hardware/software stack, from GPU architecture to application code, to drive the design of large-scale compute systems. This role involves profiling, bottleneck identification, software implementation, benchmark support, and tool development.

What you'd actually do

  1. Understand, analyze, profile, and optimize AI training workloads on state-of-the-art hardware and software platforms.
  2. Identifying performance bottlenecks of AI training on GPUs, prioritizing and then solving problems across the key AI training workloads.
  3. Implement production-quality software across multiple layers of NVIDIA's deep learning platform stack, from drivers to DL frameworks.
  4. Build and support NVIDIA submissions for MLPerf Training benchmarks.
  5. Implement key DL training workloads in NVIDIA's proprietary processor and system simulators to enable future architecture studies.

Skills

Required

  • PhD in CS, EE or CSEE (or equivalent experience) with 5+ years of relevant experience; or MS with 8+ years of experience.
  • Strong background in deep learning and neural networks, particularly in training.
  • Solid understanding of computer architecture and familiarity with GPU architecture fundamentals.
  • Proven background in analyzing and tuning application performance.
  • Proven experience with processor and system-level performance modeling.
  • Proficiency in programming with C++, Python, and CUDA.

What the JD emphasized

  • squeeze every last clock cycle
  • peak performance
  • performance analysis and optimization
  • AI training workloads
  • performance bottlenecks
  • performance modeling

Other signals

  • AI training workloads
  • performance analysis and optimization
  • GPU architecture
  • deep learning platform stack