Software Engineer, Gemini Eval Infra, Deepmind

Google Google · Big Tech · San Diego, CA +1

Software Engineer focused on building and optimizing the evaluation infrastructure for frontier AI agents and models at Google DeepMind. This role involves designing distributed execution engines, creating abstractions for LLM agent loops and tool use, implementing robust error handling and observability, and collaborating with research teams to define evaluation requirements. The goal is to provide a scalable backbone for measuring AI capabilities, safety, and readiness for launch.

What you'd actually do

  1. Design and optimize distributed evaluation execution engines capable of orchestrating large volumes of inference steps across TPU and GCU pools with high throughput and low latency.
  2. Build foundational abstractions to evaluate complex Large Language Model (LLM) agent loops, tool use, and automated LLM-as-a-judge rating systems.
  3. Design robust error classification, automated retry policies, and observability dashboards to maintain strict Service Level Objectives (SLOs) for evaluation pipeline success rates.
  4. Partner closely with GDM research scientists and Data Science teams to anticipate frontier model evaluation requirements and translate them into elegant infrastructure solutions.
  5. Mentor fellow engineers, set high standards for code quality (Python in Google3), and advocate testing and system design practices.

Skills

Required

  • software development
  • data structures and algorithms
  • distributed systems
  • software architecture
  • system design
  • Python

Nice to have

  • MBA or Master's degree in Computer Science, Software Engineering, or a related field
  • large-scale distributed systems
  • data pipelines

What the JD emphasized

  • measuring the intelligence of our prototypes
  • high-velocity, and scalable evaluation backbone
  • evaluations are the steering wheel of AI progress
  • frontier model evaluation requirements
  • large-scale distributed systems

Other signals

  • evaluating frontier AI models
  • evaluation backbone for AI development
  • measuring intelligence of AI prototypes
  • AI agents
  • large-scale distributed systems