Senior AI Systems and Algorithms Engineer

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +5 · Remote

Senior GenAI Algorithms Engineer at NVIDIA focused on advancing foundation model development, training, and deployment. The role involves large-scale distributed training, reinforcement learning, model efficiency, multimodal AI, and open-source AI infrastructure, covering the entire GenAI lifecycle from data preparation to inference optimization.

What you'd actually do

  1. Design scalable systems for preparing high-quality multimodal datasets for frontier foundation model training.
  2. Develop algorithms and systems that improve the scalability, efficiency, and cost of large-scale pre-training and post-training.
  3. Advance techniques that improve inference performance, reduce deployment cost, and enable efficient serving across cloud and edge platforms.
  4. Develop reusable infrastructure and contribute brand new model support to NVIDIA's open-source GenAI training platform.

Skills

Required

  • MS or Ph.D in Computer Science, AI, Applied Mathematics, or a related field (or equivalent experience)
  • 5+ years of relevant industry experience
  • Strong foundation in machine learning, deep learning, and optimization
  • Excellent software engineering skills, including Python and PyTorch
  • Experience building high-performance software for large-scale AI systems
  • Strong analytical, debugging, and performance optimization skills
  • Excellent communication and collaboration skills

Nice to have

  • Distributed training at scale, including Megatron-LM, Megatron Bridge, FSDP, TP/PP/CP/DP, heterogeneous or per-module parallelism, optimizer research, and efficient sparse or long-context attention.
  • Supervised fine-tuning (SFT), reinforcement learning for LLMs (e.g., PPO, GRPO, asynchronous RL), and large-scale RL frameworks such as NeMo-RL.
  • Model compression techniques including quantization (FP8, NVFP4, INT4), pruning, knowledge distillation, neural architecture search, and diffusion or non-autoregressive language models.
  • Contributing to open-source AI frameworks such as Megatron-LM, Megatron Bridge, NeMo-RL, or Hugging Face Transformers along with experience in GPU performance optimization, distributed systems, latency/throughput analysis, and profiling of large-scale AI workloads.

What the JD emphasized

  • large-scale distributed training
  • reinforcement learning for LLMs/VLMs
  • model efficiency
  • multimodal AI
  • open-source AI infrastructure
  • Megatron-LM
  • Megatron Bridge
  • NeMo-RL
  • large-scale pre-training
  • post-training
  • inference performance
  • deployment cost
  • efficient serving
  • cloud and edge platforms
  • open-source AI stack
  • Megatron-LM
  • Megatron Bridge
  • NeMo-RL
  • Hugging Face Transformers
  • GPU performance optimization
  • distributed systems
  • latency/throughput analysis
  • profiling of large-scale AI workloads

Other signals

  • large-scale distributed training
  • reinforcement learning for LLMs/VLMs
  • model efficiency
  • multimodal AI
  • open-source AI infrastructure