Senior Deep Learning Scientist, Multimodal Agentic RL

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +1 · Remote

NVIDIA is seeking a Senior Deep Learning Scientist to advance streaming and agentic multimodal AI. The role involves developing models for reasoning, planning, and acting across diverse modalities, using the Nemotron platform. Responsibilities include applying research to train and deploy neural networks for agentic systems, advancing post-training and alignment methods, researching agentic reasoning and perception, and leading multimodal dataset development and evaluation.

What you'd actually do

  1. Apply fundamental and applied research to develop, train, fine-tune, and deploy advanced neural networks for language processing in agentic systems encompassing audio-visual reasoning, tool usage, and document understanding.
  2. Advance post-training and alignment methods including instruction tuning, preference optimization, and RLHF/RLVR to improve multimodal agents for complex use cases.
  3. Research and develop agentic reasoning and grounded perception capabilities, focusing on planning, tool execution, and long-horizon task completion across digital and physical environments.
  4. Lead the collection, development, and benchmarking of multimodal datasets, ensuring high-quality evaluation of model accuracy, safety, and task completion success.

Skills

Required

  • Python
  • PyTorch
  • Transformers
  • mixture-of-experts models
  • reinforcement learning algorithms
  • post-training multimodal models
  • dataset versioning
  • experiment tracking
  • evaluation pipelines

Nice to have

  • publication record in top-tier AI and machine learning venues
  • training and deploying multimodal foundation models using large-scale distributed infrastructure
  • deep reinforcement learning techniques to train multimodal agents in complex simulation or gaming environments
  • building embodied AI systems that integrate multimodal perception with backend action-fulfillment and long-horizon planning

What the JD emphasized

  • Master’s degree (or equivalent experience) or PhD in Computer Science, AI, or Applied Math with 8+ years of relevant work experience.
  • Excellent programming skills in Python with strong fundamentals in scalable model development and deep learning frameworks like PyTorch.
  • Strong knowledge of ML/DL techniques and modern foundation model architectures, including Transformers and mixture-of-experts models.
  • Foundational understanding of reinforcement learning algorithms and implementation, including MDPs, policies, and reward design.
  • Hands-on experience in post-training multimodal models for audio-visual reasoning and human-AI interaction.
  • Proven ability to manage model development life cycles, including dataset versioning, experiment tracking, and evaluation pipelines.

Other signals

  • agentic multimodal AI
  • foundational models
  • Nemotron platform
  • LLM
  • multimodal AI products