Senior Staff Software Engineer, Tpu Performance

Google Google · Big Tech · Sunnyvale, CA +1

Senior Staff Software Engineer focused on optimizing the performance of ML training and inference workloads on Google's Tensor Processing Units (TPUs). This role involves developing and maintaining benchmarks, identifying performance opportunities through compiler/runtime optimizations, and participating in algorithmic innovations for new TPU hardware features. The role also touches on post-training optimizations like quantization and LoRA.

What you'd actually do

  1. Identify and maintain ML training and serving benchmarks that are representative to Google production and the broader ML industry.
  2. Achieve performance for customer launches, and in case of third-party/OSS models, for engaged benchmark submissions (ML Commons, InferenceX, etc).
  3. Use the benchmarks to identify performance opportunities and drive both near-term state of the art (e.g. custom kernels) and out-of the box performance (compiler/runtime optimizations, agentic tooling, auto-sharding) directly and in collaboration with partner teams.
  4. Participate in algorithmic innovations exploiting new TPU hardware features and model-preserving optimizations (speculative decoding, sparsity, quantization, LoRA, etc).
  5. Participate in co-designing models that are TPU-friendly to showcase model quality at performance excellent to OSS models typically designed on GPUs.

Skills

Required

  • Python
  • C++
  • ML infrastructure
  • model deployment
  • model evaluation
  • data processing
  • debugging
  • fine tuning
  • Speech/audio
  • reinforcement learning
  • design and architecture
  • software product testing and launching

Nice to have

  • PyTorch
  • JAX
  • performance analysis
  • performance debugging
  • cross-host systems performance
  • technical leadership
  • cross-functional project experience
  • Master’s degree or PhD

What the JD emphasized

  • 8 years of experience in software development
  • 7 years of experience leading technical project strategy, ML design, and working with ML infrastructure
  • 5 years of experience with one or more of the following: Speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.
  • 5 years of experience with design and architecture; and testing/launching software products.

Other signals

  • TPU performance optimization
  • ML training and inference
  • compiler and runtime optimizations
  • model-preserving optimizations