Research Scientist, Takeoff Intel

Anthropic Anthropic · AI Frontier · San Francisco, CA · AI Research & Engineering

Research Scientist focused on measuring and understanding recursive-self-improvement in AI systems. This role involves designing evaluations, building quantitative models of capability growth, running experiments, and assessing AI R&D acceleration. The output is focused on graded assessments and system-card sections rather than traditional publications.

What you'd actually do

  1. Identify the signals that track AI R&D acceleration and design the evaluations that measure them
  2. Build quantitative models of capability growth and self-improvement dynamics, grounded in evaluation and telemetry data
  3. Run experiments and evals to test hypotheses about automation and capability
  4. Make opinionated research bets and own the outcome
  5. Write graded assessments of what our measurements show, for internal decision-makers and public reporting

Skills

Required

  • hands-on research on large models (pretraining, fine-tuning, RL, evals, or agents scaffolds)
  • strong quantitative instincts
  • comfortable with quantitative modeling and reasoning
  • experience in forecasting
  • design an evaluation from a vague question and defend the methodology
  • write clearly and calibrate
  • motivated by impact
  • care about AI safety

Nice to have

  • Trained or RL'd frontier models hands-on
  • Experience with scaling laws, capability forecasting, or emergent-capability studies
  • A physics, applied-math, or similarly quantitative background that moved into ML
  • Written a system card section, capability report, or methodology document that others cite
  • Experience supervising and correcting AI-written code

What the JD emphasized

  • measuring and understanding recursive-self-improvement
  • design the evaluations that measure them
  • quantitative models of capability growth
  • test hypotheses about automation and capability
  • graded assessments

Other signals

  • measuring recursive-self-improvement
  • designing evaluations for AI R&D acceleration
  • quantitative modeling of capability growth
  • testing hypotheses about automation and capability