Senior Performance Co-design Engineer, LLM Training

Google Google · Big Tech · Sunnyvale, CA +1

Senior Performance Co-design Engineer focused on LLM training workloads to optimize Google's custom AI silicon (TPUs). Analyzes distributed training bottlenecks and influences future hardware design by working with ML research and compiler teams.

What you'd actually do

  1. Lead hardware/software co-design and performance modeling for LLM training workloads across current and future generation silicon.
  2. Analyze distributed training bottlenecks (compute, memory, networking) for massively scaled 1P and 3P models.
  3. Build and enhance modeling infrastructure, simulators, and performance tooling tailored to distributed ML training.
  4. Drive data-backed decisions that influence the roadmap for future TPU/Cloud Silicon architectures.
  5. Work closely with ML research and compiler teams to optimize training algorithms and influence future hardware design.

Skills

Required

  • performance modeling
  • computer architecture
  • hardware/software co-design
  • C++
  • Python

Nice to have

  • distributed ML training methodologies
  • data parallelism
  • tensor parallelism
  • pipeline parallelism

What the JD emphasized

  • LLM training workloads
  • future generation silicon
  • distributed training bottlenecks
  • future TPU/Cloud Silicon architectures
  • future hardware design

Other signals

  • LLM training workloads
  • Google's custom ML accelerators
  • TPU Chip Architecture
  • performance modeling
  • distributed training bottlenecks