Software Engineer Iii, AI and Infrastructure, Gvisor

Google Google · Big Tech · Sunnyvale, CA +2

Software Engineer III role focused on gVisor, an application kernel for containers, to enhance performance and compatibility for AI/ML workloads, particularly sandboxed hardware-accelerated inference and agentic workloads. The role involves increasing gVisor performance, supporting new workload types, collaborating on agentic workload sandboxing standards, and contributing to the open-source project.

What you'd actually do

  1. Increase gVisor performance and compatibility with new types of workloads(including implementing any syscalls or upstream Linux features that may be required). Specifically, sandboxed hardware-accelerated inference tasks and computer-use workflows will be important.
  2. Work with the GKE team to shape the industry standard for running agentic workloads and sandboxing (Collaborating on implementation of new testing, monitoring, and orchestration infrastructure).
  3. Develop an understanding of the overall project and current technical roadmap. Proactively define milestones, identify features or deficiencies that matter to users (current and future). Work with other engineers on the team to guide these projects.
  4. Work with team to invest and grow the OSS project. (define the OSS strategy, providing clarity and plans around Google-internal vs external features, and ensuring that business objectives are met while supporting a healthy OSS community).

Skills

Required

  • software development
  • large-scale infrastructure
  • distributed systems
  • networks
  • compute technologies
  • storage
  • hardware architecture
  • cloud compute platforms
  • Linux Kernels internals
  • Java
  • C/C++
  • Python
  • JavaScript
  • Go

Nice to have

  • network and security protocols in operating systems
  • Unix
  • Linux
  • device driver development
  • GPU drivers
  • operating system kernels
  • CPU scheduling
  • memory management
  • resource allocation
  • container runtimes
  • scheduling engines
  • Kubernetes
  • Docker
  • containerd
  • performance optimization
  • tuning
  • debugging of distributed systems
  • cloud-based applications
  • network stacks
  • network protocols
  • network device drivers

What the JD emphasized

  • accelerating AI inference performance
  • defining the standard runtime defense for enterprise AI agents
  • agentic workloads

Other signals

  • accelerating AI inference performance
  • defining the standard runtime defense for enterprise AI agents
  • sandboxing at scale
  • securing internal Google workloads