Staff Software Engineer, Ai/ml, Google Distributed Cloud, Storage

Google Google · Big Tech · Kirkland, WA +1

Staff Software Engineer on the AI Storage team within Google Distributed Cloud (GDC). This role focuses on building the foundational data layer for next-generation AI/ML and generative AI workloads, specifically optimizing storage systems for high-performance, massive throughput, and ultra-low latency required for AI training and inference at the edge and in air-gapped environments. The engineer will design, develop, and maintain scalable, distributed storage solutions, drive technical execution with external partners, and ensure security and availability.

What you'd actually do

  1. Design and develop scalable, distributed File and Object storage solutions that serve as the critical foundational backbone for complex AI/ML workloads within the GDC environment.
  2. Engineer advanced solutions optimized for massive throughput, specifically enabling high-frequency model checkpointing and the ultra-low latency data access required for AI training and inference.
  3. Drive technical execution with external partners (e.g., VAST Data) to seamlessly integrate industry-leading, high-performance storage hardware with Google’s distributed software ecosystem.
  4. Take full life-cycle responsibility for core storage services, ensuring uncompromising security, data durability, and high availability across disconnected edge and on-premises data center environments.
  5. Partner closely with AI infrastructure, compute, and networking teams to architect system-wide improvements, eliminate Input/Output bottlenecks, and deliver a unified "cloud-anywhere" experience.

Skills

Required

  • Software development
  • ML design
  • ML infrastructure
  • model deployment
  • model evaluation
  • data processing
  • debugging
  • fine tuning
  • systems architecture
  • distributed systems
  • large-scale storage architectures
  • systems programming (C++, Go, or Rust)

Nice to have

  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field
  • Data structures and algorithms
  • Technical leadership
  • Cross-functional project experience

What the JD emphasized

  • 8 years of experience in software development
  • 5 years of experience with ML design and ML infrastructure
  • 5 years of experience in systems architecture
  • 5 years of experience in systems programming (e.g., C++, Go, or Rust)

Other signals

  • building foundational data layer for AI/ML workloads
  • optimizing data pipeline to keep GPUs saturated
  • delivering high-performance, massively scalable storage systems for AI