Deep Learning Software Engineer, Inference - New College Grad 2026

NVIDIA NVIDIA · Semiconductors · CA +4 · Remote

Software Engineer specializing in Deep Learning Inference, focusing on designing, building, and optimizing GPU-accelerated software for efficient large-scale model serving and inference. The role involves contributing to high-performance open-source frameworks, improving platforms for LLM and Generative AI models, and optimizing performance across NVIDIA accelerators using tools like CUTLASS, OAI Triton, NCCL, and CUDA kernels.

What you'd actually do

  1. Performance optimization, analysis, and tuning of DL models in various domains like LLM, Multimodal and Generative AI.
  2. Scale performance of DL models across different architectures and types of NVIDIA accelerators.
  3. Contribute features and code to NVIDIA’s inference libraries, vLLM and SGLang, FlashInfer and LLM software solutions.
  4. Work with cross-collaborative teams across frameworks, NVIDIA libraries and inference optimization innovative solutions.

Skills

Required

  • Pursuing or recently completed a MS or PhD Computer Engineering, Computer Science, EECS, AI or related field or equivalent experience.
  • Software development experience.
  • Excellent C/C++ programming and software design skills.

Nice to have

  • SW Agile skills are helpful
  • Python experience is a plus
  • Prior experience with training, deploying or optimizing the inference of DL models in production is a plus
  • Prior background with performance modeling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus
  • GPU programming experience (CUDA, OAI TRITON or CUTLASS) is a plus
  • Contribute to deep learning software projects, such as PyTorch, vLLM, and SGLang to drive advancements in the field.
  • Experience with Multi GPU Communications (NCCL, NVSHMEM)

What the JD emphasized

  • GPU programming experience (CUDA, OAI TRITON or CUTLASS) is a plus
  • Prior experience with training, deploying or optimizing the inference of DL models in production is a plus
  • Prior background with performance modeling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus

Other signals

  • GPU-accelerated software
  • high-performance open-source frameworks
  • efficient large-scale model serving and inference
  • performance improvements for LLM and Generative AI models
  • NVIDIA accelerators
  • open-source tools and plugins
  • model serving pipelines