Senior Software Engineer, Machine Learning, Acceleration Platform

Google Google · Big Tech · Singapore

Senior Software Engineer on the Acceleration Platform team, architecting next-generation AI-native agentic systems to eliminate developer toil. Focuses on defining paradigms for autonomous AI in complex engineering challenges, establishing AI engineering excellence and security-by-design, and scaling technical capabilities through mentorship. The role involves leading the design of scalable, fault-tolerant multi-agent systems, defining best practices for LLM orchestration and RAG pipelines, establishing AI quality and safety strategies with automated evaluation frameworks, managing complex edge cases with advanced telemetry, and driving technical alignment and mentorship.

What you'd actually do

  1. Lead the design and architecture of highly scalable, fault-tolerant systems where multi-agent networks reason, plan, and execute complex workflows across vast, distributed codebases.
  2. Define best practices for the team and broader organization. Blend traditional distributed systems architecture with advanced LLM orchestration, complex Retrieval Augmented Generation (RAG) pipelines, and optimization.
  3. Establish the overarching technical strategy for AI quality and safety. Build automated evaluation frameworks that measure performance, enforce strict security standards, and reliably mitigate at scale.
  4. Manage the most intricate non-deterministic edge cases. Build advanced telemetry and introspection tooling that allows the entire organization to understand, debug, and optimize self-sustaining behavior.
  5. Drive technical alignment across local pods and global organizations. Mentor junior and mid-level engineers, translate extreme ambiguity into actionable technical roadmaps, and shape the future of AI-driven developer productivity.

Skills

Required

  • software programming in Python or C++
  • software design and architecture
  • generative AI
  • Natural Language Processing (NLP)
  • computer vision
  • speech/audio
  • reinforcement learning
  • recommendation systems
  • ML infrastructure
  • model training
  • model inference
  • model deployment
  • model evaluation
  • optimization
  • data processing
  • debugging

Nice to have

  • data structures and algorithms
  • technical leadership
  • distributed systems architecture
  • LLM orchestration
  • RAG systems
  • autonomous agentic workflows
  • core ML concepts
  • attention
  • embeddings
  • encoders/decoders
  • PyTorch
  • JAX
  • TensorFlow
  • AI safety
  • enterprise security
  • advanced prompt engineering
  • scalable model evaluation methodologies

What the JD emphasized

  • AI-native agentic systems
  • autonomous AI resolves complex engineering challenges
  • AI engineering excellence
  • security-by-design
  • multi-agent networks reason, plan, and execute complex workflows
  • LLM orchestration
  • Retrieval Augmented Generation (RAG) pipelines
  • AI quality and safety
  • automated evaluation frameworks
  • strict security standards
  • intricate non-deterministic edge cases
  • advanced telemetry and introspection tooling
  • AI-driven developer productivity

Other signals

  • AI-native agentic systems
  • eliminate systemic developer toil
  • autonomous AI resolves complex engineering challenges
  • high-ambiguity, zero-to-one initiatives
  • AI engineering excellence
  • scaling the technical capabilities of the entire organization