Lead, Hardware Deployment Engineer

xAI xAI · AI Frontier · Memphis, TN · Data Center

Lead the end-to-end bring-up of GPU compute hardware for AI training clusters, building and leading a team responsible for integration, bring-up, and repair. Drive aggressive timelines, develop processes, and manage vendors to ensure maximum hardware availability and deployment velocity.

What you'd actually do

  1. Lead, hire, and develop a dedicated hardware deployment team (deployment engineers, deployment technicians, and repair technicians) with full ownership of team structure and staffing.
  2. Own L11 rack integration and compute hardware bring-up across multiple data halls concurrently, from delivery dock to healthy production handoff.
  3. Drive aggressive bring-up timelines: achieve 95%+ node availability within days of rack delivery and 100% closure within one week per data hall.
  4. Own post-L11 hardware health: run systematic health pushes to sustain greater than 98% node availability prior to turnover to operations.
  5. Internalize non-RMA hardware repairs to maximize hardware recovery, minimize repair backlogs, and reduce dependence on OEM turnaround times.

Skills

Required

  • Leading technician or engineering teams
  • Deploying, integrating, or repairing compute/server hardware at data center scale
  • L11 (rack-level) integration and bring-up of GPU or accelerator-based systems
  • Troubleshooting servers, GPUs, NVLink/fabric interconnects, high-speed networking, and liquid cooling systems

Nice to have

  • NVIDIA GB200/GB300 NVL72 or similar rack-scale liquid-cooled GPU systems
  • Standing up a new team or function
  • Managing OEM/ODM vendor relationships
  • Hardware failure analysis, RMA processes, and component-level repair strategies
  • Data center automation, burn-in/validation tooling, and hardware health telemetry
  • Driving step-change improvements in deployment velocity or cost

What the JD emphasized

  • end-to-end bring-up of GPU compute hardware across the world's largest AI training clusters
  • critical path activities in the company
  • deployment velocity
  • institutionalize processes
  • 5+ years of hands-on experience deploying, integrating, or repairing compute/server hardware at data center scale
  • Direct experience with L11 (rack-level) integration and bring-up of GPU or accelerator-based systems
  • Demonstrated experience leading technician or engineering teams in a fast-paced deployment, manufacturing, or data center environment
  • Deep troubleshooting skills across servers, GPUs, NVLink/fabric interconnects, high-speed networking, and liquid cooling systems