Customer Engineer, AI Infrastructure, Google Public Sector

Google Google · Big Tech · Reston, VA +2

Customer Engineer focused on AI Infrastructure and HPC for Google Public Sector. Partners with technical sales to understand customer business issues, troubleshoot, conduct PoCs, and architect cloud solutions using Google AI Infrastructure and HPC ecosystem. Provides domain expertise on hardware accelerators, ML frameworks, and model building techniques, making recommendations on hardware, framework selection, and benchmarks. Manages customer relationships and collaborates with internal teams.

What you'd actually do

  1. Accelerate customer time-to-value on the largest AI Infrastructure and High Performance Computing (HPC) workloads in Google Public Sector.
  2. Build a trusted advisory relationship with customer architects, engineering leadership, and research teams. Identify customer priorities, technical objections and design strategies focused on Google AI Infrastructure and HPC ecosystem to deliver business value and resolve blockers.
  3. Provide domain expertise around hardware accelerators (GPU/TPU), prevailing ML Frameworks (PyTorch, Keras, JAX), and model building techniques.
  4. Make recommendations on Graphics Processing Unit/Tensor Processing Unit (GPU/TPU) hardware, framework selection, benchmarks, and model building required to successfully implement a complete solution.
  5. Manage the holistic research engineering relationship with customers by collaborating with specialists, product management, technical teams, and more.

Skills

Required

  • cloud native architecture in a customer-facing or support role
  • engaging with, or presenting to, technical stakeholders or executive leaders
  • machine learning (ML) model development and deployment
  • Active US Government Top Secret/Sensitive Compartmentalized Information (TS/SCI) security clearance with polygraph
  • travel up to 20% of the time

Nice to have

  • prevailing ML development frameworks (e.g., Keras, PyTorch, Tensorflow, JAX)
  • GPU and TPU based infrastructure
  • prevailing AI related tooling (Slurm, vLLM, Ray, Vertex, K8s, etc.)
  • AI software development life cycle (data processing, model building, training, evaluation, deployment)
  • deliver results and work cross-functionally to position and orchestrate a solution consisting of multiple products

What the JD emphasized

  • Active US Government Top Secret/Sensitive Compartmentalized Information (TS/SCI) security clearance with polygraph.

Other signals

  • customer-facing
  • technical sales
  • AI infrastructure
  • HPC
  • GPU/TPU
  • ML frameworks