Director, Software Engineering, AI Data Enrichment Pipelines

Google Google · Big Tech · San Jose, CA +1

This role leads a large engineering organization (~120 people) responsible for AI data enrichment pipelines that support foundational model training for Google's core products like Gemini and Search. The focus is on building and scaling platforms for data ingestion, preparation, and quality control, including multimodal data and post-training techniques like RLHF and fine-tuning.

What you'd actually do

  1. Define the strategy and lead the consolidation of disparate training data pipelines (text, HTML, images, and video) into a cohesive modularized and multimodal platform, enabling multimodal data journeys.
  2. Drive technical innovations in dataset management representation, semantic data filtering, and post-training quality loops (e.g., Reinforcement Learning from Human Feedback (RLHF), fine-tuning data prep) where industry standards are constantly shifting.
  3. Serve as the primary technical liaison engaging with senior leadership across GDM, Search, and AI Foundations to align roadmaps and secure resource commitments.
  4. Lead the evolution of information retrieval infrastructure toward improved usability, reliability, scalability, efficiency, architectural clarity, and coherence with a focus on AI model development journeys.

Skills

Required

  • Bachelor's degree in Computer Science, Electrical Engineering, a related technical field or equivalent practical experience.
  • 15 years of experience in software engineering, with 8 years in a technical engineering, systems architecture, or principal-level role.
  • 10 years of experience leading and managing software engineering organizations.
  • Experience managing other engineering managers.

Nice to have

  • Master’s degree or PhD in Computer Science, Machine Learning, Computer Engineering, or related technical field.
  • Deep, specialized technical knowledge of large-scale distributed systems, planet-scale data processing, and complex data orchestration workflows.
  • Strong understanding of machine learning lifecycle pipelines, specifically data preparation requirements for post-training, fine-tuning, and reinforcement learning (RLHF).
  • Demonstrated ability to engage, influence, and collaborate with senior executive stakeholders (VPs, Directors) across multiple distinct business units.
  • Proven track record of leading large, complex, cross-functional engineering organizations (100+ people) in a high-growth, fast-paced environment.
  • Proven track record of technical innovation within planet-scale distributed systems and data pipelines.

What the JD emphasized

  • expert-level grasp of technical tradeoffs
  • lead a high-performing engineering organization of ~120 people
  • managing other engineering managers
  • Deep, specialized technical knowledge of large-scale distributed systems, planet-scale data processing, and complex data orchestration workflows.
  • Proven track record of leading large, complex, cross-functional engineering organizations (100+ people)

Other signals

  • foundational models
  • training data
  • planet-scale data processing
  • multimodal data journeys
  • post-training quality loops