Staff+ Software Engineer, Capacity Engineering

Anthropic Anthropic · AI Frontier · San Francisco, CA · Software Engineering - Infrastructure

Staff+ Software Engineer for Capacity Engineering at Anthropic. This role focuses on building and operating production systems for managing and optimizing a large-scale AI infrastructure fleet. Responsibilities include developing data pipelines for telemetry ingestion, creating observability tooling for fleet health, and implementing performance instrumentation to measure hardware utilization across training, inference, and evaluation workloads. The role involves working with cloud providers, Kubernetes, Python, and SQL, and requires a strong understanding of systems engineering and data engineering principles to ensure efficient resource allocation and cost attribution.

What you'd actually do

  1. Build the planning and allocation stack — the tools leadership uses to allocate capacity, teams use to plan against their allocations, and the scheduler enforces. Cross-region and cross-provider placement, guardrails, queueing, occupancy KPIs.
  2. Drive the efficiency programs: stranding and rightsizing, unused capacity recovery, and job-level utilization across training, inference, and eval. Establish per-config baselines and work with system-owning teams to close the gaps. At this fleet size a single point of utilization is worth eight figures a month.
  3. Own attribution and forecasting — reconcile billing across ten-plus providers against telemetry and internal systems, attribute spend to the workloads that generate it, and turn demand signals and research roadmaps into a defensible compute plan and supply pipeline.
  4. Build the data platform underneath all of it: pipelines ingesting occupancy, utilization, and cost from a rapidly diversifying fleet into BigQuery, with real ownership of completeness, latency SLOs, and gap detection. Every new provider is a net-new integration.
  5. Operate Kubernetes-native systems at scale — collection agents, workload labeling, and the taint/reservation/scheduling behavior that determines what capacity is actually usable.

Skills

Required

  • Python
  • SQL
  • major cloud provider (Amazon Web Services, Google Cloud, or Microsoft Azure)
  • observability tooling stack (Prometheus, PromQL, Grafana)
  • Kubernetes

Nice to have

  • data engineering
  • systems engineering

What the JD emphasized

  • production-quality code
  • production systems
  • production quality
  • production quality