Senior Production Engineer, Managed Cloud

Crusoe Crusoe · Data AI · San Francisco, CA - US · Cloud Engineering

This role focuses on the production engineering of Crusoe's AI-optimized cloud platform, specifically ensuring the reliability and scalability of managed AI services for LLM workloads. The engineer will design, operate, and optimize large-scale training and inference clusters, focusing on performance, reliability, and cost-efficiency. The role involves defining and measuring SLIs/SLOs, automating observability, and contributing to the architecture of AI-first distributed systems.

What you'd actually do

  1. Design and operate reliable managed AI services with a focus on serving and scaling LLM workloads
  2. Define, measure, and improve SLIs/SLOs across to ensure performance and reliability targets are met
  3. Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters
  4. Automate observability by building telemetry and performance tuning strategies for latency-sensitive services
  5. Contribute to the architecture of next-generation distributed systems purpose-built for AI-first environments

Skills

Required

  • Strong software engineering background
  • experience building production-grade systems
  • Demonstrated experience in distributed systems design and implementation
  • SRE mindset and experience
  • Defining and measuring SLIs/SLOs
  • Building monitoring and observability systems
  • Driving performance and reliability improvements
  • Designing fault-tolerant systems and automated testing strategies
  • Proficiency in at least one modern programming language (Python, Go, Java, C++)
  • Familiarity with Kubernetes or container orchestration platforms

Nice to have

  • collaboration and communication skills
  • Ability to thrive in a fast-paced, mission-driven environment

What the JD emphasized

  • serving and scaling LLM workloads
  • latency-sensitive services
  • large-scale training and inference clusters
  • performance and reliability targets

Other signals

  • AI infrastructure
  • LLM workloads
  • training and inference clusters
  • latency-sensitive services