Forward Deployed Engineer, Genai, Google Cloud, Data

Google Google · Big Tech · Singapore

This role focuses on deploying and engineering complex AI applications, specifically GenAI, within enterprise environments on Google Cloud. The Forward Deployed Engineer will work with massive datasets, build production-grade agentic workflows, design data pipelines for evaluation datasets, and construct ETL/ELT pipelines for semantic modeling. The role emphasizes transitioning prototypes to production, ensuring data quality, and enabling AI-driven analytics like Text-to-SQL.

What you'd actually do

  1. Serve as a developer for complex AI applications, transitioning from rapid prototypes to production-grade agentic workflows that drive measurable Return on Investment (ROI).
  2. Co-build with customer engineering teams to instill Google-grade development best practices, ensuring long-term project success and high end-user adoption.
  3. Design and build high-throughput batch and streaming data pipelines and utilities to curate multi-terabyte evaluation datasets and execute offline/online evaluation generation jobs for model intelligence mining.
  4. Construct scalable ETL/ELT pipelines using Dataform, dbt, BigQuery, or Dataproc to design enterprise data marts and semantic modeling layers specifically engineered to maximize data quality, schema clarity, and accuracy for Text-to-SQL and natural language analytical interfaces.
  5. Create mechanisms for large-scale synthetic data generation that maintain strict referential integrity across complex relational schemas, leveraging advanced tools and custom generative utilities for privacy-safe model benchmarking and fine-tuning.

Skills

Required

  • software development
  • data engineering
  • SQL
  • Python
  • Java
  • Scala
  • Go
  • ETL/ELT frameworks
  • dbt
  • Dataform
  • enterprise data modeling layers
  • data marts

Nice to have

  • Master's degree or PhD in Computer Science, Data Science, Artificial Intelligence
  • integrating semantic metadata formats
  • enterprise taxonomies
  • ontologies
  • large-scale data warehouses and lakes
  • designing batch, offline, and online evaluation harnesses
  • intelligence mining jobs
  • benchmarking LLM capabilities
  • Text-to-SQL accuracy
  • semantic parsing
  • configuring and deploying secure code execution harnesses
  • interpreter sandboxes
  • automated data analysis
  • synthetic data generation at scale
  • multi-table referential integrity
  • Faker
  • Snowfakery
  • custom constraint engines

What the JD emphasized

  • production-grade agentic workflows
  • high-throughput batch and streaming data pipelines
  • multi-terabyte evaluation datasets
  • enterprise data marts and semantic modeling layers
  • large-scale synthetic data generation
  • strict referential integrity

Other signals

  • Deploying AI/ML models into production environments
  • Building complex data pipelines for AI applications
  • Working with large-scale enterprise data