Staff/principal Devops Engineer, AI Inference

Lila Sciences Lila Sciences · AI Frontier · Alewife, Cambridge, MA · Software

Staff/Principal DevOps Engineer - AI Inference role focused on designing, implementing, and optimizing infrastructure for serving machine learning models at scale. This involves building low-latency, high-throughput inference systems on GPU clusters and cloud accelerators, collaborating with ML engineers and scientists to ensure reliable production deployment and maximize compute efficiency.

What you'd actually do

  1. GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads
  2. Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing
  3. Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency
  4. Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads
  5. Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments

Skills

Required

  • DevOps
  • SRE
  • Platform Engineering
  • GPU/accelerator infrastructure
  • Kubernetes for ML workloads
  • AWS
  • Terraform
  • Helm
  • Model serving infrastructure
  • Networking for distributed inference
  • Python

Nice to have

  • LLM inference optimization
  • Continuous batching
  • Speculative decoding
  • Quantization
  • Tensor parallelism
  • Pipeline parallelism
  • Multiple accelerator families
  • Multi-region deployment
  • Rust
  • Go
  • Chaos engineering on GPU workloads
  • Incident management
  • Capacity modeling
  • Model registries
  • Artifact versioning
  • ML supply chain security
  • Observability platform expertise
  • Startup/high-growth experience

What the JD emphasized

  • Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale
  • Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management
  • Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)
  • Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks
  • Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7
  • Strong proficiency in Python for automation, tooling, and integration with ML frameworks

Other signals

  • GPU/accelerator infrastructure on Kubernetes
  • Model serving platforms
  • Intelligent request routing and load balancing
  • Autoscaling systems
  • Production-grade deployment pipelines for ML models
  • Infrastructure-as-code
  • Observability and performance optimization
  • CI/CD pipelines for model artifacts
  • AWS cloud infrastructure for ML
  • Cost optimization and capacity planning