Senior Software Engineer, Gemini Audio, Deepmind

Google Google · Big Tech · Mountain View, CA +2

Senior Software Engineer on the Gemini Audio team at Google DeepMind, focusing on research and development of audio representations for Gemini models. The role involves designing and building training infrastructure, optimizing for TPUs using XLA, implementing tools for analysis, and collaborating on audio components and inference efficiency. Requires strong software development skills, experience with ML optimization, and familiarity with accelerators and ML frameworks.

What you'd actually do

  1. Design, build, and maintain training infrastructure to support Gemini Audio encoder and pretraining models, focusing on Accelerated Linear Algebra (XLA) optimization for training on Tensor Processing Units (TPUs).
  2. Implement tools to analyze and track audio model training efficiency and health.
  3. Contribute and work closely with Gemini to build and maintain audio related components.
  4. Collaborate with research teams across Gemini Audio to improve the model training and inference efficiency.

Skills

Required

  • software development in Python
  • software design and architecture
  • compiler optimization
  • code generation
  • runtime systems for popular accelerators
  • Machine Learning
  • Machine Learning Optimization
  • Performance Optimization
  • Large Language Models

Nice to have

  • TPU architectures
  • memory hierarchies
  • performance bottlenecks
  • Accelerated Linear Algebra
  • tailoring algorithms and ML models to exploit TPU architecture strengths and minimize weaknesses
  • compilers (OpenXLA, MLIR)
  • serving libraries/frameworks (vLLM, sglang)
  • ML frameworks (JAX, PyTorch)
  • debugging experience to improve performance of single-mode or multi-mode (distributed) systems

What the JD emphasized

  • Accelerated Linear Algebra (XLA) optimization
  • Tensor Processing Units (TPUs)
  • Machine Learning, Machine Learning Optimization, Performance Optimization, and Large Language Models
  • TPU architectures, memory hierarchies, performance bottlenecks, and Accelerated Linear Algebra
  • tailoring algorithms and ML models to exploit TPU architecture strengths and minimize weaknesses
  • compilers (OpenXLA, MLIR), serving libraries/frameworks (vLLM, sglang) and ML frameworks (JAX, PyTorch)

Other signals

  • Gemini Audio
  • audio representations
  • Gemini models
  • audio-to-audio dialog models
  • text-to-speech
  • speech-to-speech translation
  • duplex audio dialogs
  • training infrastructure
  • encoder and pretraining models
  • Accelerated Linear Algebra (XLA) optimization
  • Tensor Processing Units (TPUs)
  • audio model training efficiency and health
  • model training and inference efficiency