Staff Software Engineer, Code RL

Anthropic Anthropic · AI Frontier · San Francisco, CA · AI Research & Engineering

Staff Software Engineer focused on the engineering aspects of reinforcement learning for AI coding capabilities, specifically creating and scaling agentic coding environments. The role involves designing frameworks and APIs for researchers, managing production RL runs, and improving the reliability and structure of research codebases. It emphasizes Python expertise, API design, and anticipating system failures.

What you'd actually do

  1. Design widely-used APIs, frameworks, and abstractions that other engineers and researchers build on, with careful attention to interface legibility and principled defaults
  2. Embed with research teams on a rotational basis: understand their engineering needs, build systems and APIs that support their work, and transfer ownership so teams can maintain those systems after you rotate off
  3. Work directly in research codebases, improving reliability and structure without slowing down the research they support
  4. Anticipate silent failure modes and prevent them structurally through type safety, well-designed invariants, targeted testing, and refactors that shrink the surface area for bugs
  5. Contribute to the reliability of production RL systems, including monitoring, regression detection, and triage tooling

Skills

Required

  • Python
  • static typing
  • safe async and concurrency patterns
  • performant Python code
  • API design
  • framework design
  • system design
  • type safety
  • testing
  • written and verbal communication
  • ambiguity tolerance

Nice to have

  • infrastructure for ML research
  • tooling for ML research
  • frameworks for RL workflows
  • reinforcement learning concepts
  • agentic systems
  • LLM training pipelines
  • large-scale distributed systems
  • client libraries
  • SDKs
  • sandboxed execution platforms
  • containerized execution platforms
  • remote execution platforms
  • large-scale data processing
  • dataset lifecycle management
  • plugin systems
  • extensible class hierarchies
  • embedding with research teams
  • consulting for other teams
  • handing off systems
  • code standards
  • lint rules
  • static verification approaches
  • technical lead experience
  • setting engineering standards
  • open source project maintenance

What the JD emphasized

  • Deep expertise in Python
  • track record of designing intuitive, safe APIs or frameworks
  • Experience working productively in large, evolving, or research-style codebases
  • Demonstrated ability to anticipate failure modes — especially silent ones — and prevent them structurally
  • Comfortable diving into messy research code
  • hard-won intuition for how complex systems fail — especially silently

Other signals

  • Reinforcement learning
  • Agentic coding environments
  • Scaling RL efforts
  • Production RL runs
  • Sandboxed execution