Lead AI Validation and Test Engineer

AMD AMD · Semiconductors · Austin, TX · Engineering

Lead AI Validation and Test Engineer at AMD responsible for system and silicon validation of AMD EPYC Server & AMD Instinct products, focusing on AI rack validation. The role involves defining and driving validation strategy for AI/ML server platforms, including CPU, GPU, memory, firmware, and software, with an emphasis on system debug, automation, and data-driven methodologies to ensure product quality and readiness for customer deployments in hyperscale environments.

What you'd actually do

  1. Lead system-level, rack-level, and cluster-scale validation strategy for AI and machine learning server platforms.
  2. Define and drive comprehensive validation plans covering customer deployment scenarios across scale-up and scale-out environments.
  3. Develop validation methodologies and test strategies spanning CPU, GPU, memory, BIOS, BMC, networking, storage, platform firmware, operating systems, and infrastructure components.
  4. Collaborate with architecture, hardware, firmware, software, and platform engineering teams to define validation requirements, identify coverage gaps, and improve overall product quality.
  5. Lead investigation and root cause analysis of complex hardware, firmware, software, networking, and system integration issues.

Skills

Required

  • System and Silicon validation
  • Post-silicon validation
  • System debug
  • Validation strategy
  • Debug methodology improvements
  • AI Rack Validation
  • Scalable and automated solutions
  • Technical leadership
  • Validation and debug expertise
  • Product development, definition, root cause and resolution
  • Agility and collaborative approach
  • Engineering teams and other stakeholders (System Architects, IP design, SoC, FW, SW, manufacturing)
  • CPU, GPU, memory, BIOS, BMC, networking, storage, platform firmware, operating systems, and infrastructure components
  • Hardware, firmware, software, networking, and system integration issues
  • Scalable validation frameworks, automation infrastructure, telemetry-driven workflows, and data-driven validation methodologies
  • Large-scale validation data analysis
  • Validation readiness criteria, quality metrics, and coverage strategies
  • Mentoring and technical leadership
  • Validation methodologies, automation capabilities, debugging effectiveness, and system-level engineering practices
  • Execution across multiple validation programs
  • Managing priorities, dependencies, schedules, and technical risks
  • Validation status, readiness assessments, quality metrics, technical recommendations, and key risks
  • Customer-focused validation efforts
  • Test environments accurately represent real-world deployment conditions and hyperscale operating environments
  • Continuous improvements in validation processes, tools, and automation
  • OS, FW, Silicon, and HW issues debugging
  • Industry standard busses and their software stack
  • X86 architecture, SoC design, memory, RAS & power management
  • System architecture, technical debug, and validation strategy
  • Platform/ system level debug, Operating System, Device Drivers and System BIOS interactions
  • Datacenter industry technologies and their software stack
  • Self-starter, able to independently drive tasks to completion

Nice to have

  • Bachelors or Masters degree in electrical or computer engineering

What the JD emphasized

  • Lead system-level, rack-level, and cluster-scale validation strategy for AI and machine learning server platforms.
  • Define and drive comprehensive validation plans covering customer deployment scenarios across scale-up and scale-out environments.
  • Develop validation methodologies and test strategies spanning CPU, GPU, memory, BIOS, BMC, networking, storage, platform firmware, operating systems, and infrastructure components.
  • Collaborate with architecture, hardware, firmware, software, and platform engineering teams to define validation requirements, identify coverage gaps, and improve overall product quality.
  • Lead investigation and root cause analysis of complex hardware, firmware, software, networking, and system integration issues.

Other signals

  • AI server platforms
  • AI and machine learning server platforms
  • customer deployment scenarios
  • hyperscale operating environments
  • large-scale validation data