Staff Software Engineer, ML Systems Co-design

Google Google · Big Tech · Sunnyvale, CA +2

Staff Software Engineer focused on ML Systems Co-Design, working at the intersection of ML workloads and TPU silicon roadmaps. The role involves defining strategy for open-source ML models on TPUs, analyzing mapping gaps, and optimizing agentic workflows ahead of silicon launch. Requires deep technical proficiency in ML performance modeling, compilers, and ML execution frameworks.

What you'd actually do

  1. Deliver certified performance ladders 6 months pre-pilot as the definitive goals for pricing, capacity planning, and XLA/kernel optimization.
  2. Adapt emerging open-source models into TPU-native paired variants (e.g., sparsity optimization) to prove asymmetric TPU superiority.
  3. Ingest and optimize multi-turn agentic workflows, reasoning loops, and prompt/decode disaggregated topologies ahead of silicon.
  4. Author high-signal reference implementations (PyTorch/JAX/Pallas) that guarantee reachability of pre-pilot goals under real-world compiler constraints.
  5. Partner with director and executive technical leadership across TPU Hardware, XLA Compiler, and Cloud AI Systems to steer multi-year roadmaps.

Skills

Required

  • C++
  • Python
  • software testing
  • software launching
  • performance analysis
  • large-scale systems data analysis
  • visualization tools
  • debugging
  • software design
  • software architecture

Nice to have

  • ML workload performance modeling on pre-silicon systems
  • high level compiler intermediate representations (e.g., MLIR dialects, XLA, HLO)
  • ML execution frameworks (JAX, PyTorch, PyTorch/XLA)
  • accelerator kernel programming (Pallas, Triton, CUDA)
  • architecting, validating, or extending high-performance simulation tools, emulation frameworks, or analytical performance modeling pipelines (e.g., roofline models, cycle-accurate or compiler-aware simulators)
  • L6-level ability to lead complex, multi-quarter technical initiatives across distinct organizational boundaries (e.g., hardware design, compilers, AI research, and cloud infrastructure)

What the JD emphasized

  • ML workload performance modeling
  • high level compiler intermediate representations
  • ML execution frameworks
  • accelerator kernel programming
  • architecting, validating, or extending high-performance simulation tools, emulation frameworks, or analytical performance modeling pipelines
  • L6-level ability to lead complex, multi-quarter technical initiatives

Other signals

  • ML Systems Co-Design
  • TPU silicon roadmaps
  • open-source ML models
  • performance analysis
  • ML workload performance modeling