AI Infrastructure Engineer, Serving Platform

Scale AI Scale AI · Data AI · London, United Kingdom · EPD

ML Infrastructure Engineer focused on building and maintaining scalable, reliable, and efficient serving platforms for LLMs and other models. The role involves backend system design, integrating models, and developing monitoring solutions, bridging research and engineering.

What you'd actually do

  1. Build and maintain fault-tolerant, high-performance systems for serving LLMs and other models at scale.
  2. Build an internal platform to empower LLM capability discovery.
  3. Collaborate with researchers and engineers to integrate and optimize models for production and research use cases.
  4. Conduct architecture and design reviews to uphold best practices in system design and scalability.
  5. Develop monitoring and observability solutions to ensure system health and performance.

Skills

Required

  • Python
  • Go
  • Rust
  • C++
  • LLM serving
  • LLM routing
  • rate limiting
  • token streaming
  • load balancing
  • budgets
  • reasoning
  • tool calling
  • prompt templates
  • Docker
  • Kubernetes
  • AWS
  • GCP
  • Terraform
  • backend system design
  • scalability
  • monitoring
  • observability

Nice to have

  • vLLM
  • SGLang
  • TensorRT-LLM
  • text-generation-inference

What the JD emphasized

  • 4+ years of experience building large-scale, high-performance backend systems
  • Experience with LLM serving and routing fundamentals
  • Experience with LLM capabilities and concepts

Other signals

  • building platforms for scalable, reliable, and efficient serving of LLMs
  • ML infrastructure
  • backend system design
  • integrating and optimizing models for production