Software Engineering Manager, Google Cloud Tpu

Google Google · Big Tech · Tel Aviv, Israel

Software Engineering Manager for Google Cloud TPU, leading an engineering team focused on critical TPU software services for AI customers. The role involves architecting and executing large-scale ML infrastructure, understanding LLM operations from chip to fleet, solving complex ML infrastructure issues, driving technical strategy, fostering team culture, and working with customers.

What you'd actually do

  1. Build, mentor, and inspire a software engineering team. Cultivate a culture of high performance, innovation, psychological safety, and deliberate career development where every engineer grows.
  2. Shape the technical strategy and multi-year product roadmap, seamlessly aligning team objectives with the broader Google Cloud goal and our pioneering AI-first initiative.
  3. Drive engineering velocity across the entire project lifecycle, ensuring the predictable, high-quality delivery of scalable features.
  4. Drive deep technical decision-making for our infrastructure components, fostering a culture of code reviews, scalable architecture design, AI-first and data-driven development, while mentoring engineers.
  5. Advocate an unyielding customer-first mindset, establishing standards for service reliability, performance, and operational excellence, transforming complex customer requirements into technical realities, and maintaining a healthy production.

Skills

Required

  • software development
  • building distributed cloud services
  • engineering management
  • guiding software engineering teams
  • hiring
  • team development
  • LLM training or inference
  • performance optimizations
  • distributed execution
  • GPU or TPU acceleration
  • PyTorch
  • JAX
  • TensorFlow

Nice to have

  • Master's degree or PhD in Computer Science or related technical field
  • working in a complex, matrixed organization

What the JD emphasized

  • critical Google Cloud TPU software services
  • large-scale ML infrastructure
  • LLM operations from chip level to fleet level
  • complex ML infrastructure issues
  • AI-first initiative
  • LLM training or inference
  • performance optimizations
  • distributed execution
  • GPU or TPU acceleration

Other signals

  • leading ML infrastructure
  • LLM operations from chip to fleet
  • customer AI solutions