Engineering Manager, Foundation Model Inference (fmapi)

Databricks Databricks · Data AI · Mountain View, CA · Engineering

Engineering Manager for Databricks' Foundation Model Inference team, focusing on building and scaling the infrastructure for serving, scaling, and optimizing frontier models with enterprise-grade reliability and performance. The role involves leading a team of infrastructure engineers, shaping product roadmaps for real-time, provisioned throughput, and batch inference, and ensuring operational excellence for large-scale inference workloads.

What you'd actually do

  1. Lead and grow a team of talented, product-minded infrastructure Engineers working on Databricks Foundation Model Inference and Foundation Model API; The teams owns the systems powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama).
  2. Help shape the roadmap for products spanning real-time inference, provisioned throughput, and batch inference use cases.
  3. Partner closely with product and engineering leadership to deliver the right capabilities with strong reliability, quality, and service health.
  4. Build an inclusive, high-performing team that attracts, develops, and retains exceptional engineers.
  5. Maintain a deep understanding of your team’s technical area and uphold a strong bar for architecture, implementation quality, and operational excellence.

Skills

Required

  • Experience managing and growing high-performing software engineering teams.
  • Strong technical judgment in distributed systems, platform infrastructure, AI/ML infrastructure, or large-scale backend services, with the ability to maintain quality even in areas you have not worked on personally.
  • A track record of delivering complex, multi-quarter engineering initiatives with high quality and predictable execution.
  • Experience partnering effectively with product management and peer engineering teams to translate customer and business needs into roadmaps and shipped outcomes.
  • Operational rigor, including experience with service health, incident response, postmortems, and continuous improvement.
  • Excellent communication, coaching, and hiring skills.
  • BS in Computer Science or related field.

What the JD emphasized

  • large-scale inference workloads
  • enterprise production workloads
  • large multi-quarter initiatives
  • high quality and predictable execution
  • operational rigor

Other signals

  • build and operate the premier data and AI infrastructure platform
  • enables our customers to leverage deep data insights for transformative business improvements
  • build strong teams, uphold a high technical bar, and drive execution on large multi-quarter initiatives in a fast-moving AI infrastructure environment
  • own the systems powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama)