Sr Technical Program Manager, Aws Generative AI & ML Servers

Amazon Amazon · Big Tech · Seattle, WA · Project/Program/Product Management--Technical

The Senior Technical Program Manager will drive the end-to-end delivery of GPU-accelerated servers powering AI/ML workloads across AWS's global fleet. This role involves facilitating requirements gathering, coordinating cross-functional engineering teams (hardware, firmware, software, operations), managing ODM partnerships, and establishing quality feedback systems. The position requires technical depth to challenge engineering workstreams and program management skills to deliver complex hardware at global scale, ensuring reliability and cost-effectiveness for AI/ML infrastructure.

What you'd actually do

  1. Facilitate requirements gathering sessions with internal customers and stakeholders.
  2. Drive cross-functional alignment across hardware, firmware, software, and operations teams to maintain development schedules.
  3. Identify program risks by challenging technical workstreams and connecting technical constraints to schedule impact.
  4. Define acceptance criteria and coordinate qualification testing across teams.
  5. Conduct knowledge transfer sessions with receiving teams.

Skills

Required

  • 5+ years of technical product or program management experience
  • 7+ years of working directly with engineering teams experience
  • Experience managing programs across cross functional teams, building processes and metrics
  • Experience with hardware development lifecycles
  • Experience with supply chain and manufacturing processes
  • Experience with cloud infrastructure
  • Experience with AI/ML infrastructure

Nice to have

  • Experience with GPU-accelerated servers
  • Experience with Nvidia GPU offerings
  • Experience with AWS services
  • Experience with firmware development
  • Experience with security and safety standards

What the JD emphasized

  • end-to-end delivery
  • GPU-accelerated servers
  • AI/ML workloads
  • global fleet
  • cross-functional engineering teams
  • hardware and software disciplines
  • ODM partnerships
  • multiple continents
  • closed-loop quality feedback systems
  • operational data
  • design improvements
  • technical depth
  • program management excellence
  • complex hardware
  • global scale
  • program schedules
  • identifying blockers early
  • escalate dependencies
  • critical path
  • AI/ML infrastructure
  • reliability commitments
  • data-driven quality systems
  • fleet-wide hardware failures
  • corrective actions
  • operational metrics
  • feedback loops
  • operational telemetry
  • validation criteria
  • future designs
  • mission-critical AI/ML workloads

Other signals

  • GPU-accelerated servers
  • AI/ML workloads
  • hardware and software systems
  • program management excellence
  • global scale