Tech Lead Manager, Business Agent Quality

Google Google · Big Tech · Mountain View, CA +1

Tech Lead Manager for Business Agent Quality at Google, focusing on evaluating, observing, and scaling AI agent performance across Google's commerce ecosystem. The role involves leading the development of evaluation infrastructure and automated rating systems, bridging applied AI and engineering to create a self-service platform, and solving reliability issues for LLM reasoning and agent pipelines. This includes mentoring an engineering team.

What you'd actually do

  1. Own the technical goal for evaluating, observing, and iterating on AI agent performance across Google’s commerce ecosystem.
  2. Lead the development of evaluation infrastructure and automated rating systems (AutoRaters) to measure conversation helpfulness and factuality at scale.
  3. Bridge applied AI and engineering excellence to transition evaluation frameworks into a self-service platform for internal partners.
  4. Solve complex reliability issues to maintain high quality standards for LLM reasoning and low-latency agent pipelines.
  5. Provide direction while remaining to mentor and grow a nimble, high-performing engineering team.

Skills

Required

  • software engineering and architecture experience
  • testing and launching software products
  • integrating generative AI tools or LLM interfaces into workflows
  • establishing standards for evaluation and quality metrics
  • people manager experience

Nice to have

  • data structures and algorithms
  • working in a complex, matrixed organization
  • Generative AI experience
  • conversational AI reliability metrics
  • agent orchestration harnesses with validation protocols
  • collaborating with Product, Data Science, and user-facing feature teams

What the JD emphasized

  • evaluating AI agent performance
  • scaling AI agent performance
  • evaluation infrastructure
  • automated rating systems
  • LLM reliability
  • agent pipelines

Other signals

  • evaluating AI agent performance
  • scaling LLM reliability
  • building evaluation infrastructure
  • scaling AI agent performance