Senior Performance Co-design Engineer, LLM Serving

Google Google · Big Tech · Sunnyvale, CA +1

Senior Performance Co-Design Engineer focused on optimizing LLM serving performance on custom AI silicon (TPUs). The role involves analyzing and optimizing inference latency and throughput, developing simulation and profiling tools, and co-designing architectural improvements with hardware and software teams to influence future accelerator roadmaps.

What you'd actually do

  1. Conduct comprehensive serving performance studies on (current and emerging) LLMs (1P/3P models).
  2. Develop and maintain advanced simulation, profiling, and modeling tools to identify bottlenecks, understand key characteristics and project serving workload performance.
  3. Partner with model researchers, software and hardware teams to co-design architectural improvements tailored to Large Language Model (LLM) inference latency and throughput.
  4. Drive data-backed decisions that influence the roadmap for future TPU/Cloud Silicon architectures.

Skills

Required

  • performance modeling/engineering
  • computer architecture
  • co-design
  • systems engineering
  • C++
  • Python

Nice to have

  • hardware/software co-design
  • pre-silicon performance analysis
  • large-scale ML model optimization
  • ML infrastructure
  • profiling tools
  • deep learning inference/serving optimizations
  • accelerator architectures

What the JD emphasized

  • LLM Serving
  • inference latency and throughput
  • performance modeling/engineering
  • co-design
  • ML models

Other signals

  • LLM Serving
  • TPU Chip Architecture
  • Performance Co-design
  • Inference Latency and Throughput
  • ML Accelerators