Evals Infrastructure Tech Lead / Manager

Anthropic Anthropic · AI Frontier · San Francisco, CA · AI Research & Engineering

Lead the team building and scaling the distributed systems that orchestrate, schedule, and execute evals for frontier models, ensuring measurement quality, reproducibility, and that eval signal reaches decision-makers. This role involves managing engineers and contributing directly as an engineer, focusing on inference, research, and infrastructure engineering.

What you'd actually do

  1. Lead the team building the distributed systems that schedule, orchestrate, and execute evals for our frontier model training
  2. Own eval throughput and cost: compute allocation across suites, queueing against constrained accelerator pools, caching and reuse of eval work
  3. Build and scale the harnesses researchers use to define, run, and iterate on evals
  4. Make eval results trustworthy — determinism, reproducibility, and honest uncertainty quantification on reported metrics
  5. Ensure eval signal reaches the dashboards and reviews where launch decisions actually get made

Skills

Required

  • Python
  • Rust
  • leading technical projects end-to-end on large-scale distributed systems
  • high-throughput, fault-tolerant systems on cloud or on-prem accelerator fleets
  • communicating with researchers and translating research needs into infrastructure

Nice to have

  • LLM inference or training infrastructure
  • eval or benchmarking systems, especially agentic evals requiring sandboxed execution
  • working statistical literacy — variance, confidence intervals, sample-size sufficiency for noisy metrics
  • observability and regression detection over time-series metrics

What the JD emphasized

  • 1+ years managing engineers
  • measurement quality, not just pipeline uptime
  • eval or benchmarking systems
  • agentic evals requiring sandboxed execution
  • observability and regression detection over time-series metrics

Other signals

  • evals infrastructure
  • large scale distributed systems
  • frontier models
  • measurement quality
  • launch decisions