Senior Deep Learning Software Engineer, Inference and Model Optimization

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +1 · Remote

Senior Deep Learning Software Engineer focused on optimizing generative AI models (LLMs, diffusion) for inference efficiency using techniques like pruning, sparsity, quantization, and automated deployment. The role involves applied research and developing a software platform (TRT Model Optimizer), working with frameworks like PyTorch, HuggingFace, and high-performance kernels in CUDA, TRT-LLM, and Triton.

What you'd actually do

  1. Train, develop, and deploy state-of-the generative AI models like LLMs and diffusion models using NVIDIA's AI software stack.
  2. Leverage and build upon the torch 2.0 ecosystem (TorchDynamo, torch.export, torch.compile, etc...) to analyze and extract standardized model graph representation from arbitrary torch models for our automated deployment solution.
  3. Develop high-performance optimization techniques for inference, such as automated model sharding techniques (e.g. tensor parallelism, sequence parallelism), efficient attention kernels with kv-caching, and more.
  4. Collaborate with teams across NVIDIA to use performant kernel implementations within our automated deployment solution.
  5. Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities.

Skills

Required

  • Python
  • PyTorch
  • HuggingFace
  • Deep Learning
  • software design
  • debugging
  • performance analysis
  • test design
  • algorithms
  • programming fundamentals
  • communication skills

Nice to have

  • Contributions to PyTorch, JAX, or other Machine Learning Frameworks
  • GPU architecture
  • compilation stack
  • end-to-end performance debugging
  • TensorRT
  • CUDA
  • CUTLASS
  • Triton

What the JD emphasized

  • 5+ years of relevant work or research experience in Deep Learning
  • Excellent software design skills, including debugging, performance analysis, and test design
  • Strong proficiency in Python, PyTorch, and related ML tools (e.g. HuggingFace)
  • Strong algorithms and programming fundamentals

Other signals

  • optimizing generative AI models
  • inference efficiency
  • automated deployment strategies
  • high-performance kernel implementations