Backend Staff Software Engineer, Consumer Colab, Cloud AI

Google Google · Big Tech · Kirkland, WA +1

This role focuses on building and maintaining infrastructure for AI-centric runtimes and serving infrastructure to support AI/ML workflows and agent workloads. It involves developing and maintaining the largest fleet of Jupyter runtimes globally, with a focus on scalability, performance, and supporting new offerings like headless Colab with agent-friendly CLI, public API, and hosted MCP. The role also involves developing novel abuse detection and ensuring support for modern AI/agent workloads.

What you'd actually do

  1. Develop and maintain infrastructure to access data science runtimes to power runtimes as an AI tool compatible with common agents and works well with our infra - payments, auth, security, abuse, public api, cli, mcp etc.
  2. Contribute to maintaining the largest fleet of data science runtimes in the world - keeping our ecosystem up to date, our machine fleet up to date and lowering toil to support the fleet.
  3. Collaborate with the broader Colab teams to design and implement software systems and tools that support the cutting edge AI/ML workflows.
  4. Contribute to our Tier2 oncall.
  5. Communicate effectively across engineering and product teams.

Skills

Required

  • C++
  • Golang
  • Rust
  • software design
  • architecture
  • large-scale infrastructure
  • distributed systems
  • networks
  • compute technologies
  • storage
  • hardware architecture
  • testing
  • launching software products

Nice to have

  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field
  • data structures
  • algorithms
  • technical leadership
  • complex, matrixed organization
  • cross-functional projects
  • notebook products

What the JD emphasized

  • largest fleet of Jupyter runtimes
  • AI-centric runtimes
  • serving infrastructure
  • agent friendly CLI
  • public API
  • modern AI/agent workloads
  • data science runtimes
  • AI/ML workflows

Other signals

  • AI-centric runtimes
  • serving infrastructure
  • agent friendly CLI
  • public API
  • modern AI/agent workloads
  • data science runtimes
  • AI/ML workflows