Staff Software Engineer, Environments Infrastructure

Anthropic Anthropic · AI Frontier · San Francisco, CA · AI Research & Engineering

Staff Software Engineer focused on building and maintaining the infrastructure for reinforcement learning environments at Anthropic. This role involves designing APIs, frameworks, and tooling to support researchers in building and operating AI agents, with a focus on stateful distributed systems, agent runtimes, and productionizing research for AI capabilities improvement.

What you'd actually do

  1. Design widely used APIs, frameworks, and abstractions that other engineers and researchers build on, making correct usage the default and ruling out entire classes of errors structurally
  2. Own the platform layers that sit beneath every environment, including the agent runtime
  3. Build the tooling that lets environment owners understand, debug, and maintain their environments in production without needing an infrastructure engineer in the loop
  4. Embed with research teams on a rotational basis, work directly in their codebases without slowing down the research they support, and transfer ownership when you rotate off
  5. Anticipate silent failure modes and prevent them structurally through type safety, well-designed invariants, targeted testing, and refactors that reduce the room for correctness issues

Skills

Required

  • Deep expertise in Python, including static typing, safe async and concurrency patterns, and writing performant code
  • Strong taste in API and framework design, the ability to explain why an interface is right or wrong rather than just recognizing it, and a track record of other engineers or teams adopting and building on frameworks you have built
  • Experience designing or operating stateful concurrent or distributed systems, and reasoning carefully about failure, retires, idempotency, and consistency
  • A habit of verification: you measure before you conclude, and you build the checks that let a system show it's correct
  • Experience working productively in large, evolving, or research-style codebases that you didn't originally write
  • Strong written and verbal communication with collaborators of varied engineering backgrounds, and comfort with ambiguity: able to scope your own work from a loosely defined problem and drive it to a maintainable outcome

Nice to have

  • Experience building infrastructure, tooling, or frameworks for machine learning research or RL workflows, and familiarity with agentic systems or LLM training pipelines
  • Experience building agent frameworks, orchestration engines, or multi-agent systems, including checkpoint and restore, replay, and coordination of long-running stateful processes
  • Experience using AI coding tools on code where correctness matters, with good judgment about what to delegate and how to make the results verifiable
  • Experience building client libraries or SDKs on top of sandboxed, containerized, or remote execution platforms
  • Experience with large-scale data processing, dataset lifecycle management, or data lineage systems
  • Experience designing serialization schemes, plugin systems, or extensible class hierarchies used across an organization
  • Experience embedding with or consulting for other teams and handing off systems for others to own, or defining code standards adopted across teams, or prior experience as a technical lead

What the JD emphasized

  • stateful distributed system
  • agent runtime
  • production RL runs
  • state-sharing and recovery model
  • multi-agent workloads
  • sandboxed execution platform
  • failure and retry model
  • agentic systems
  • LLM training pipelines

Other signals

  • productionize research
  • frameworks and APIs
  • agent runtime
  • production RL runs
  • stateful distributed system
  • workflow engine
  • actor framework
  • durable-execution runtime
  • messy research code
  • abstractions
  • RL environment abstraction
  • model-tool interface
  • state-sharing and recovery model
  • multi-agent workloads
  • sandboxed execution platform
  • failure and retry model
  • agentic systems
  • LLM training pipelines