Technical Program Manager, RL Research

Anthropic Anthropic · AI Frontier · San Francisco, CA · Technical Program Management

Technical Program Manager supporting Reinforcement Learning research, focusing on managing the research lifecycle from pre-training through post-training. The role involves tracking experiments, driving research reviews, establishing processes, and collaborating with various teams to identify blockers and make technical trade-off decisions. Requires a background in ML engineering or research, with hands-on experience in ML training pipelines, RLHF systems, and large-scale data infrastructure.

What you'd actually do

  1. Deliver a regular read on the ground truth in RL research, covering performance against baselines, experiment results, day-to-day health, and incidents
  2. Work with RL org leads on prioritizing, ranking, and tracking the state of experiments
  3. Drive research reviews end to end in partnership with set the agenda, make sure the right context is in the room ahead of time, and close the loop on what gets decided
  4. Establish processes and frameworks that bring structure to an unstructured research setting without slowing researchers down
  5. Collaborate with research leads, infrastructure engineers, and data operations to identify blockers, prioritize competing needs, and make technical trade-off decisions

Skills

Required

  • ML engineering or ML research background
  • ML training pipelines
  • RLHF systems
  • large-scale data infrastructure
  • execution plans
  • process development
  • stakeholder management
  • communication skills

Nice to have

  • ability to debug data pipelines
  • read RL transcripts to spot issues
  • make allocation and quality decisions in real time
  • navigate a fast-growing organization
  • identify critical people and teams across research, infrastructure, product, and data operations
  • coordinate across teams without losing velocity
  • contribute meaningfully to discussions with researchers
  • influence senior technical staff

What the JD emphasized

  • ML engineering or ML research background
  • deep, hands-on experience with ML training pipelines, RLHF systems, and large-scale data infrastructure in production
  • track record of building execution plans and inventing high-leverage processes
  • fast learner who builds deep contextual understanding in unfamiliar technical domains
  • resourceful, high-agency, and able to navigate ambiguity and shifting priorities

Other signals

  • RL research
  • model development lifecycle
  • research reviews
  • execution plans
  • high-leverage processes