Director/sr. Manager, AI Inference Model Scaling

Cerebras Cerebras · Semiconductors · US and Canada Offices · Software

This role leads the Inference Model Scaling team, responsible for enabling state-of-the-art foundation models and generative AI workloads to run efficiently on Cerebras' Wafer-Scale Engine (WSE). The team builds the compiler frontend, model transformation pipeline, graph optimization infrastructure, high-performance kernel enablement, and runtime integration. The leader will define the technical vision, organizational strategy, and execution roadmap for a globally distributed engineering team, focusing on ML model compilation, optimization, and high-performance kernel development. This role requires deep technical leadership in compiler infrastructure, ML systems, and infrastructure software, along with experience leading engineering teams and driving production-quality software delivery.

What you'd actually do

  1. Define the technical roadmap and strategy for the team.
  2. Hire, mentor, and grow a high-performing engineering team.
  3. Partner with Cloud Platform, ML, and Hardware teams in planning and delivering for end-to-end service enablement in Cloud and On-Premise settings
  4. Own planning, prioritization, and execution across multiple concurrent initiatives.

Skills

Required

  • BS, MS, or PhD in Computer Science, Computer Engineering or related field.
  • 12+ years building compiler, ML systems, or infrastructure software.
  • 5+ years leading engineering teams.
  • Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).
  • Strong understanding of graph compilation and optimization.
  • Experience with Python and C++.
  • Experience delivering production-quality software.
  • Strong communication and cross-functional leadership skills.

Nice to have

  • Experience building compiler frontends for AI accelerators.
  • Experience supporting PyTorch, JAX, TensorFlow, or ONNX.
  • Experience with LLM inference or training systems.
  • Familiarity with distributed compilation.
  • Experience working with hardware architects.
  • Experience leading teams through rapid growth.

What the JD emphasized

  • 12+ years building compiler, ML systems, or infrastructure software.
  • 5+ years leading engineering teams.
  • Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).
  • Strong understanding of graph compilation and optimization.
  • Experience delivering production-quality software.

Other signals

  • Enabling state-of-the-art foundation models and generative AI workloads to run efficiently on Cerebras' Wafer-Scale Engine (WSE).
  • Build the compiler frontend, model transformation pipeline, graph optimization infrastructure, high-performance kernel enablement, and runtime integration.
  • Drive support for emerging LLM architectures and inference workloads.
  • Balance rapid model support with long-term ML Compiler architecture.