Tech Lead Manager, Staff Software Engineering, Xprof

Google Google · Big Tech · Sunnyvale, CA +1

Tech Lead Manager for Staff Software Engineering focused on optimizing machine learning software/hardware stacks, including accelerators and inference, within Google's AI and Infrastructure team. The role involves technical leadership, team management, and driving continuous improvements through performance debugging and analysis of ML paradigms.

What you'd actually do

  1. Be focused around driving continuous improvements to the machine learning software/hardware stacks through providing insightful performance debugging. Provide insights by summarizing different views of captured profile data such as trace timelines, memory usage, HLO profiles, ML graph summaries.
  2. Learn and build an intuitive understanding of existing data collection, analysis, and visualization workflows.
  3. Support new and exciting ML paradigms (such as horizontal scaling for upcoming TPU chips) by making contributions across the end to end stack and analysis tools.
  4. Partner with product area leads to understand model optimization use cases, drive cross functional efforts to deliver on chip profiling requirements, and propose new hardware features.
  5. Collaborate across Hardware, Driver, Runtime, and Performance Analysis teams and many other stakeholders.

Skills

Required

  • software development
  • programming
  • debugging
  • computer architecture
  • C++
  • technical leadership
  • people management

Nice to have

  • Python
  • standard ML
  • high-performance computing
  • embedded systems
  • machine learning
  • accelerator architectures
  • driver/run times

What the JD emphasized

  • technical leadership
  • manage engineers
  • people management
  • technical leadership role
  • performance debugging
  • ML paradigms
  • hardware/software interactions with accelerators
  • chip profiling requirements

Other signals

  • machine learning paradigms for training/inference
  • hardware/software interactions with accelerators
  • performance debugging
  • ML graph summaries
  • horizontal scaling for upcoming TPU chips
  • chip profiling requirements