Forward Deployed Engineer, Data, Genai, Google Cloud

Google Google · Big Tech · Singapore

This role focuses on deploying and optimizing AI applications, specifically GenAI, within large enterprise customer environments using Google Cloud's data and AI stack. The engineer will work with complex data lakes and warehouses to build production-grade workflows, including Text-to-SQL, automated evaluations, and data intelligence solutions. Key responsibilities involve data pipeline engineering, semantic data modeling, evaluation harness setup, and ensuring long-term reliability and adoption. The role emphasizes transitioning prototypes to production-grade agentic workflows and generating synthetic data for benchmarking and fine-tuning.

What you'd actually do

  1. Serve as a developer for complex AI applications, transitioning from rapid prototypes to production-grade agentic workflows that drive measurable Return on Investment (ROI).
  2. Co-build with customer engineering teams to instill Google-grade development best practices, ensuring long-term project success and high end-user adoption.
  3. Design and build high-throughput batch and streaming data pipelines and utilities to curate multi-terabyte evaluation datasets and execute offline/online evaluation generation jobs for model intelligence mining.
  4. Construct scalable ETL/ELT pipelines using Dataform, dbt, BigQuery, or Dataproc to design enterprise data marts and semantic modeling layers specifically engineered to maximize data quality, schema clarity, and accuracy for Text-to-SQL and natural language analytical interfaces.
  5. Create mechanisms for large-scale synthetic data generation that maintain strict referential integrity across complex relational schemas, leveraging advanced tools and custom generative utilities for privacy-safe model benchmarking and fine-tuning.

Skills

Required

  • Software development
  • Data engineering
  • SQL
  • Python
  • Java
  • Scala
  • Go
  • ETL/ELT frameworks (e.g., dbt, Dataform)
  • Enterprise data modeling
  • Data marts

Nice to have

  • Master's degree or PhD in Computer Science, Data Science, Artificial Intelligence
  • Integrating semantic metadata formats, enterprise taxonomies, or ontologies
  • Designing batch, offline, and online evaluation harnesses
  • Synthetic data generation at scale
  • Configuring and deploying secure code execution harnesses and interpreter sandboxes

What the JD emphasized

  • production-grade agentic workflows
  • high-throughput batch and streaming data pipelines
  • enterprise data marts and semantic modeling layers
  • large-scale synthetic data generation

Other signals

  • Deploying AI/ML models into production environments
  • Working with large-scale enterprise data
  • Building and optimizing data pipelines for AI applications
  • Focus on customer success and ROI