Technical Lead, Google Cloud Tpu

Google Google · Big Tech · Tel Aviv, Israel

Technical Lead for Google Cloud TPU, focusing on leading the engineering team responsible for the software services that power Google's AI customers. This role involves architecting and executing large-scale ML infrastructure, with a deep understanding of LLM operations from chip to fleet levels, solving complex ML infrastructure issues, driving technical strategy, and collaborating with customers.

What you'd actually do

  1. Act as the crucial bridge between raw Tensor Processing Unit (TPU) silicon and production-ready machine learning, owning the software integration and operational ecosystem that powers Google's most advanced AI.
  2. Lead the end-to-end New Product Introduction process—coordinating complex cross-functional launches from initial concept to General Availability—while ensuring the reliability and scalability of a massive fleet of TPU chips.
  3. Drive foundational engineering efforts, such as developing the TPU runtime API, qualifying the OS images for TPU Virtual Machines (VMs) and Bare Metal instances, and managing fleet-wide reliability through advanced telemetry and automated repair workflows.
  4. Lead the architecture, technology and the overseeing of implementation of the TPU solutions to production.
  5. Analyze customers’ issues (working with customers and the field team) and translate these into viable technical solutions.

Skills

Required

  • software development
  • building distributed cloud services
  • engineering technical leadership
  • leading software engineering teams
  • LLM training
  • LLM inference
  • performance optimizations
  • distributed execution
  • GPU acceleration
  • TPU acceleration
  • PyTorch
  • JAX
  • TensorFlow
  • integrating generative AI tools
  • integrating LLM interfaces into workflows

Nice to have

  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • working in a complex, matrixed organization involving cross-functional, or cross-business projects.

What the JD emphasized

  • 8 years of experience in software development, focusing on building distributed cloud services.
  • 5 years of experience in a formal engineering technical leadership role, leading software engineering teams.
  • 2 years of experience in LLM training or inference, including performance optimizations, distributed execution, GPU or TPU acceleration, or PyTorch, JAX, or TensorFlow programming.

Other signals

  • leading architecture and execution of large-scale ML infrastructure
  • deep understanding of LLM operations from chip level to fleet levels
  • solving complex ML infrastructure issues
  • driving technical strategy
  • working directly with customers to ensure successful landings