Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma AI Luma AI · AI Frontier · SF Bay Area, CA +3 · Remote · Systems Research & Engineering

Luma AI is seeking a Research Scientist/Engineer to build and scale distributed reinforcement learning infrastructure for their multimodal foundation models. The role involves designing and implementing systems for high-throughput RL training, including environments, reward functions, and evaluation tooling, to enable models to reason, use tools, and act over long horizons.

What you'd actually do

  1. Design, build, and scale distributed RL post-training systems for large multimodal models — orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs
  2. Build and optimize high-throughput rollout generation, including efficient integration of inference engines (e.g. vLLM, SGLang) into the training loop, weight synchronization, and asynchronous / off-policy training schemes
  3. Design and implement RL environments for agentic and multi-step tasks — sandboxed code execution, tool use, computer use, and multimodal interaction — that are reproducible, hermetic, and scalable to millions of episodes
  4. Build reward infrastructure: verifiable / programmatic rewards, reward model serving, LLM-as-judge pipelines, and defenses against reward hacking
  5. Develop the evaluation, monitoring, and debugging tooling needed to keep large RL runs stable, diagnose convergence and throughput regressions, and understand model behavior mid-run

Skills

Required

  • Hands-on experience post-training LLMs with reinforcement learning (e.g. PPO / GRPO-family methods, RLHF, RLVR / RL from verifiable rewards) at meaningful scale
  • Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents — including sandboxed execution and multi-turn tool use
  • Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang)
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads

Nice to have

  • Experience running RL training across >100 GPUs, including asynchronous or disaggregated trainer/rollout architectures
  • Experience with containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads
  • Research contributions in RL for LLMs — reasoning, agents, reward modeling, or long-horizon tasks — or open-source contributions to RL training frameworks

What the JD emphasized

  • post-trained LLMs with RL
  • RL environments
  • reward functions
  • verifiers
  • evaluation harnesses
  • tool use
  • multi-turn tool use
  • RL post-training frameworks
  • inference engines
  • GPU clusters
  • networking
  • communication libraries
  • mixed training + inference workloads

Other signals

  • Reinforcement learning infrastructure for foundation models
  • Distributed training systems
  • RL environments for agentic tasks
  • Reward infrastructure and evaluation systems