Technical Program Manager, Hardware Systems

Anthropic Anthropic · AI Frontier · San Francisco, CA · Technical Program Management

This role is for a Technical Program Manager focused on hardware systems for AI model training and serving infrastructure. The individual will own the execution of hardware systems from concept to production deployment, managing ODMs, CMs, and component vendors, while coordinating with internal teams. Key responsibilities include developing program plans, driving execution with manufacturers, establishing NPI processes, managing bring-up and validation, and owning interfaces with silicon, datacenter, and software teams. The role also involves evaluating partners and technologies, communicating program status, and exploring AI-assisted approaches to hardware program management. While the role supports AI infrastructure, the core function is hardware program management, not direct AI/ML model development.

What you'd actually do

  1. Own the integrated hardware systems program plan across board, chassis, rack, and interconnect — from requirements through EVT/DVT/PVT, qualification, and production ramp — and be the source of truth on schedule, risk, and readiness.
  2. Drive execution with ODMs, contract manufacturers, and component/optics vendors: SOWs, build schedules, deliverable tracking, factory readiness, and escalation when they slip.
  3. Stand up and run the NPI cadence — build matrix, engineering reviews, DFx, issue triage, exit criteria — and keep it lightweight enough that a small, senior team actually uses it.
  4. Drive the bring-up and validation program for new platforms: lab and pilot-rack logistics, firmware/BMC/diagnostics coordination, test coverage, and the reliability and telemetry targets a platform has to hit before it enters the fleet.
  5. Own the interfaces that make or break a hardware program: silicon (samples, package, power/thermal envelopes), datacenter (power, cooling, rack integration, deployment windows), infrastructure software (management plane, provisioning, telemetry), and supply chain (long-lead components, second-sourcing, allocation).

Skills

Required

  • 8+ years in hardware systems / NPI program management for compute, networking, storage, or accelerator platforms.
  • Direct ownership of at least one platform from concept through volume production — you can walk through its EVT/DVT/PVT arc, what broke, and how you recovered.
  • Working technical fluency across the hardware stack — power delivery, thermal, mechanical, high-speed interconnect, board design, system management — sufficient to interrogate status and push back credibly on architects and vendors.
  • Experience managing ODMs/CMs and component vendors, including SOW negotiation, build tracking, quality escapes, and escalation.
  • Track record of standing up program mechanisms on a new or early-stage hardware effort, not just operating inside a mature PLC.
  • Strong written communication — plans, reviews, and exec updates that many teams depend on.
  • Comfort operating with high autonomy and limited process in a fast-moving, ambiguous environment.

Nice to have

  • Experience with AI/ML accelerator systems (GPU, TPU, custom ASIC), HPC, or hyperscale datacenter platforms and their scale-up / scale-out fabrics.
  • Experience taking a first-generation, custom-silicon-based system from bring-up through fleet deployment.
  • Experience with liquid-cooled and/or high-density rack architectures, rack level high power distribution, or optical interconnect programs.
  • Familiarity with the chip-package-system interface and coordinating tightly with a silicon team.
  • Experience with datacenter integration, deployment, and fleet-scale reliability/RAS programs.
  • Practical experience using AI tools in your own workflow, and clear ideas about where they help and where they don't.
  • BS/MS in EE, CE, ME, or a related field, or equivalent experience.

What the JD emphasized

  • shipped hardware at scale
  • realistic view of ODM/CM schedules and where they break
  • owning consequential program calls without a large organization behind them
  • at least one platform from concept through volume production
  • track record of standing up program mechanisms on a new or early-stage hardware effort