Senior Software Engineer, Fleet-level ML Performance

Google Google · Big Tech · Sunnyvale, CA +1

Senior Software Engineer focused on fleet-level ML performance analysis for TPU systems, involving hardware/software co-design, performance optimization, and defining requirements for future AI infrastructure. The role analyzes critical ML workloads like Gemini on next-generation TPUs to balance performance, cost, power, and reliability.

What you'd actually do

  1. Perform fleet-level performance analysis of key ML workloads (e.g., Gemini) on future TPU systems using advanced simulation tools to evaluate hardware/software trade-offs and guide next-generation chip architecture.
  2. Partner with teams across the ML stack including model researchers, compiler developers, systems engineers, and TPU architects to analyze and optimize performance across the design space.
  3. Design and implement a unified ML Accelerator Platform Performance Estimation Methodology using C++ and Python to enable scalable performance and TCO projections across Google.
  4. Collaborate cross-functionally with data center, hardware architecture, and framework teams to define key hardware and software requirements for future AI infrastructure.

Skills

Required

  • systems architecture
  • computer architecture
  • power and performance trade-off analysis
  • data center, cloud, infrastructure hardware optimization
  • Reliability, Availability, and Serviceability (RAS) features, paradigms, or architecture
  • C++
  • Python

Nice to have

  • deep learning workloads
  • embedding architectures
  • hardware execution characteristics

What the JD emphasized

  • critical ML workloads
  • future TPU systems
  • hardware/software trade-offs
  • next-generation chip architecture
  • ML stack
  • optimize performance
  • ML Accelerator Platform Performance Estimation Methodology
  • scalable performance
  • TCO projections
  • future AI infrastructure
  • hardware and software requirements

Other signals

  • performance analysis of critical ML workloads
  • hardware/software co-design
  • maximize TPU efficiency
  • AI infrastructure