AI and ML Infra Software Engineer, GPU Clusters - New College Grad 2026

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +1

NVIDIA is seeking an AI/ML Infrastructure Software Engineer to optimize GPU clusters for AI/ML research. The role involves collaborating with researchers to identify and resolve infrastructure issues, monitoring performance, and ensuring efficient resource utilization. The engineer will work with various teams to build a seamless AI/ML infrastructure ecosystem and stay updated on the latest AI/ML advancements.

What you'd actually do

  1. Collaborate closely with our AI and ML research teams to understand their infrastructure needs and obstacles, translating those observations into actionable improvements.
  2. Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.
  3. Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results.
  4. Collaborate with diverse teams, including researchers, data engineers, and DevOps professionals, to build a seamless and coordinated AI/ML infrastructure ecosystem.
  5. Stay on top of the latest advancements in AI/ML technologies, frameworks, and effective strategies, and promote their implementation within the company.

Skills

Required

  • AI/ML infrastructure
  • HPC workloads
  • GPU
  • storage
  • scheduling & orchestration
  • high-speed networking
  • containers technologies
  • PyTorch (DDP, FSDP)
  • NeMo
  • JAX
  • Python
  • Go
  • Bash
  • cloud computing platforms

Nice to have

  • parallel computing frameworks and paradigms

What the JD emphasized

  • proven experience in AI/ML and HPC workloads and infrastructure
  • in-depth knowledge of accelerated computing
  • Expertise in running and optimizing large-scale distributed training workloads
  • deep understanding of AI/ML workflows

Other signals

  • GPU Clusters
  • AI/ML research
  • infrastructure gaps
  • scalable solutions
  • large-scale distributed training workloads