AI Cluster Technical Program Manager – Validation, Debug & Agentic AI

AMD AMD · Semiconductors · Austin, TX · Engineering

Technical Program Manager to lead execution of AI cluster engineering programs with a focus on GPU platforms, rack-level solutions, and AI Cluster validation. This role is responsible for driving end-to-end delivery from GPU + server integration through rack bring-up, scale testing, failure analysis, and system debug closure, ensuring platform readiness for hyperscale and enterprise AI deployments. The role also involves championing Agentic AI and AIOps solutions for automated incident triage, log analysis, root-cause identification, and operational workflow automation at scale.

What you'd actually do

  1. Define, plan, and drive program plans for AI infrastructure systems validation and readiness, including server integration, rack bring-up, and cluster-scale deployment readiness.
  2. Own program execution for GPU-based AI platforms, spanning system bring-up, qualification, scale readiness, and deployment validation across server, rack, and cluster levels.
  3. Own program planning and execution for multi-node and multi-rack scale testing, including test strategy, scheduling, coverage tracking, and readiness gates.
  4. Act as the execution lead for platform debug, coordinating across engineering teams to ensure fast triage, root-cause analysis, and resolution of system-level issues.
  5. Champion Agentic AI and AIOps solutions for automated incident triage, log analysis, root-cause identification, and operational workflow automation at scale

Skills

Required

  • Program management
  • AI infrastructure
  • GPU platforms
  • System integration
  • Validation
  • Debug
  • Root cause analysis
  • Risk management
  • Incident management
  • Agentic AI
  • AIOps
  • Hardware
  • Firmware
  • Networking
  • Scale testing
  • Cross-functional team leadership
  • Executive communication

Nice to have

  • Experience with EVT/DVT/PVT phases
  • Familiarity with AI agents for automation

What the JD emphasized

  • AI infrastructure systems validation
  • GPU platforms
  • rack-level solutions
  • AI Cluster validation
  • Agentic AI capabilities
  • GPU + server integration
  • rack bring-up
  • scale testing
  • failure analysis
  • system debug closure
  • platform readiness
  • hyperscale and enterprise AI deployments
  • hardware, firmware, networking, and scale-test execution
  • GPU-based AI platforms
  • system bring-up
  • qualification
  • scale readiness
  • deployment validation
  • server, rack, and cluster levels
  • GPU, CPU, firmware, BIOS/BMC, and system teams
  • multi-node and multi-rack scale testing
  • test strategy
  • scheduling
  • coverage tracking
  • readiness gates
  • multi-rack scale testing
  • compute trays
  • switch trays
  • cabling
  • power
  • cooling
  • management infrastructure
  • rack bring-up plans
  • EVT, DVT, and scale phases
  • lab operations
  • infrastructure
  • engineering teams
  • rack access
  • power
  • networking
  • test readiness
  • scale, performance, and automation teams
  • workloads
  • stress tests
  • regressions plans
  • platform debug
  • engineering teams
  • triage
  • root-cause analysis
  • resolution of system-level issues
  • high-impact failures
  • GPU, HSIO, FW, rack, network
  • debug forums
  • ownership and closure plans
  • debug depth vs. program timelines
  • tradeoffs
  • leadership
  • risk and impact
  • system rack and cluster-level debug activities
  • validation
  • deployment
  • fleet operations
  • fault isolation
  • root-cause analysis
  • hardware, firmware, networking, and software domains
  • incident management
  • critical deployment and validation issues
  • cross-functional war rooms
  • executive communications
  • timely resolution
  • Mean Time to Detect (MTTD)
  • Mean Time to Mitigate (MTTM)
  • Mean Time to Resolution (MTTR)
  • incident recurrence rates
  • fleet health indicators
  • post-incident reviews (PIRs)
  • root-cause analysis (RCA)
  • corrective actions
  • long-term preventive measures
  • Sev1/Sev2 incident recurrence
  • fleet reliability
  • deployment readiness
  • Agentic AI and AIOps solutions
  • automated incident triage
  • log analysis
  • root-cause identification
  • operational workflow automation
  • scale
  • Agentic AI capabilities
  • validation and debug workflows
  • AI agents
  • Automated triage
  • Log analysis
  • Root-cause identification
  • Knowledge retrieval
  • Test orchestration
  • Incident management
  • data engineering teams
  • AI-powered operational tools
  • success metrics
  • measurable business impact
  • agent-based automation initiatives
  • AI-first operational workflows
  • engineering organizations
  • complex hardware or AI infrastructure programs
  • bring-up, validation, and deployment phases

Other signals

  • AI infrastructure systems validation
  • GPU platforms
  • rack-level solutions
  • AI Cluster validation
  • Agentic AI capabilities