Senior Staff Software Engineer, Ai/ml

Google Google · Big Tech · Sunnyvale, CA +1

Senior Staff Software Engineer role focused on developing and optimizing on-device AI solutions across various Google products and platforms. This involves training task-specific models, co-designing model runtimes, enabling developer testing and deployment, and improving inference performance on diverse hardware accelerators.

What you'd actually do

  1. Train task-specific models of in app AI (tiny Gemma/Juno Nano), enabling Gemma model/runtime co-design (quantization, conversion), build industry-leading on-device AI solutions (voice translate, generative image editing).
  2. Enable developers to test/evaluate/deploy across devices (Edge Portal /Model Explorer/Developer Device Platform).
  3. Develop and guide critical projects in Google's on-device ML infrastructure (e.g., LiteRT, LiteRT-LM).
  4. Enable on-device deployment of key models, such as Gemini Nano and Gemma, across various accelerators (GPU /Pixel TPU /NPUs/CPU) on Android, Chrome, and more.
  5. Improve performance of on-device model inference via optimizations in the model representation, on-device runtime and kernel implementation.

Skills

Required

  • software development
  • technical project strategy
  • ML design
  • model deployment
  • model evaluation
  • data processing
  • debugging
  • fine tuning
  • design and architecture
  • testing and launching software products
  • machine learning infrastructure
  • C++
  • performance
  • GPU programming
  • mobile GPU

Nice to have

  • on-device ML Software Development Kits (SDKs)/tooling (e.g., TensorFlow Lite, ExecuTorch, Core ML, SNPE/QNN)
  • ML converters/compilers and runtimes
  • hardware-accelerated ML inference techniques
  • generative AI model architectures
  • optimization for on-device execution
  • Android
  • iOS
  • web browsers
  • embedded devices

What the JD emphasized

  • industry-scale ML infrastructure
  • on-device deployment

Other signals

  • on-device AI
  • model deployment
  • inference optimization