Principal Engineer, Cape

Crusoe Crusoe · Data AI · San Francisco, CA - US · Cloud Engineering

Principal Engineer to build a self-driving fleet management system for AI accelerators, focusing on closed-loop autonomy, unified observability, and energy-aware compute scheduling. This involves applying ML to hardware telemetry for failure prediction and treating the fleet as a single logical computer.

What you'd actually do

  1. Make the fleet self-driving.
  2. Build a unified observability plane that correlates GPU, networking, storage, orchestration, and workload signals.
  3. Develop a closed-loop autonomy system that diagnoses, decides, and remediates issues without human intervention.
  4. Maximize goodput as an objective function by trading scheduling, placement, and maintenance decisions.
  5. Forecast hardware failures hours in advance using ML on noisy telemetry data.

Skills

Required

  • 10+ years building infrastructure-layer systems at scale
  • Deep experience with distributed systems design
  • Hands-on fluency with GPU/HPC infrastructure
  • Track record of designing and shipping large-scale observability or telemetry platforms
  • Comfort operating in ambiguity and defining architecture and standards for a system that doesn't exist yet
  • Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar)

Nice to have

  • Experience applying ML/statistical methods to noisy operational telemetry
  • Prior exposure to zero-trust or policy-based multi-tenancy architectures

What the JD emphasized

  • make the fleet self-driving
  • closed-loop autonomy
  • ML problem on noisy hardware telemetry at fleet scale
  • agentic operations
  • 0→1 charter

Other signals

  • building a self-driving fleet of AI accelerators
  • closed-loop autonomy for infrastructure
  • ML for failure prediction on hardware telemetry
  • energy-aware compute scheduling