Staff Software Engineer, Vertex Genai Evaluation Service

Google Google · Big Tech · Sunnyvale, CA +1

Staff Software Engineer focused on building and scaling evaluation frameworks and tools for GenAI models and agents within Google Cloud. This role involves developing strategies for measuring model and agent quality, designing infrastructure for evaluation, and collaborating with research teams to advance SOTA evaluation technology.

What you'd actually do

  1. Drive high-quality, scalable evaluation frameworks and tools for GenAI developers and lead features to enable seamless developer experience in each stage of GenAI evaluation.
  2. Develop and implement strategies for evaluation solutions to measure model and agent quality performance at scale, design infrastructure and product features of the evaluation frameworks to evaluate advanced agent features.
  3. Collaborate with model and agent quality teams, develop scalable systems, and ensure our evaluation workflow is efficient and provides high quality user experience.
  4. Engage with customers, investigate and understand product usability gaps, customer pain points, and lead GenAI research.

Skills

Required

  • software development
  • software design and architecture
  • ML design
  • ML infrastructure optimization
  • model deployment
  • model evaluation
  • data processing
  • debugging
  • fine tuning
  • GenAI techniques
  • LLMs
  • Multi-Modal
  • Large Vision Models
  • language modeling
  • computer vision

Nice to have

  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field
  • LLM evaluation
  • machine learning algorithms and tools
  • general generative AI development
  • complex organization involving cross-functional, or cross-business projects
  • building and maintaining scalable production systems
  • enterprise control planes
  • security horizontals

What the JD emphasized

  • 8 years of experience in software development
  • 5 years of experience testing, and launching software products
  • 3 years of experience with software design and architecture
  • 5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning)
  • 2 years of experience with GenAI techniques (e.g., LLMs, Multi-Modal, Large Vision Models) or with GenAI-related concepts (e.g., language modeling, computer vision)
  • 5 years of experience with LLM evaluation, machine learning algorithms and tools, and general generative AI development

Other signals

  • evaluation frameworks
  • model and agent evaluation
  • GenAI evaluation
  • LLM evaluation
  • quality metrics