Staff Software Engineer- Foundation Model Inference

Databricks Databricks · Data AI · Mountain View, CA · Engineering

Databricks is seeking a Staff Software Engineer for their Foundation Model Inference team. This role focuses on building and optimizing the infrastructure for serving large-scale generative AI models, ensuring enterprise-grade reliability, performance, and scalability. The engineer will work with various models and collaborate across teams to enhance AI capabilities on the Databricks platform.

What you'd actually do

  1. Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama)
  2. Improve reliability, latency, and efficiency of distributed AI workloads
  3. Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences
  4. Shape how developers and data scientists build and interact with AI on Databricks

Skills

Required

  • backend or infrastructure engineering
  • distributed systems
  • scalable APIs
  • cloud-native infrastructure
  • real-time serving
  • ML infrastructure
  • GPU orchestration
  • service-oriented architecture
  • deployment pipelines
  • system observability

Nice to have

  • SageMaker
  • Vertex AI
  • Azure ML
  • MLflow
  • PyTorch
  • Ray
  • vLLM
  • SGLang
  • developer platforms
  • internal tools supporting AI workflows

What the JD emphasized

  • enterprise-grade reliability and performance
  • scale inference workloads
  • real-time serving
  • ML infrastructure
  • GPU orchestration
  • system observability

Other signals

  • serving frontier models
  • enterprise-grade reliability and performance
  • scale inference workloads