Senior Product Manager, Orchestration

Crusoe Crusoe · Data AI · San Francisco, CA - US · Product and Design

Senior Product Manager for Orchestration at Crusoe, an AI infrastructure company. The role focuses on owning major product areas within the Managed Orchestration portfolio, including Kubernetes cluster lifecycle, control-plane capabilities, node pools, workload scheduling (Slurm on Kubernetes), and customer-facing APIs. This is a deeply technical role requiring close collaboration with engineering to define requirements, examine edge cases, and ensure product quality for demanding AI workloads. Requires strong technical fluency in Kubernetes, cloud infrastructure, distributed systems, and experience with AI/ML infrastructure interactions.

What you'd actually do

  1. Own Major Orchestration Product Areas: Take end-to-end ownership of substantial product surfaces, including their roadmap, outcomes, priorities, quality, and continued improvement.
  2. Develop Durable Technical Product Designs: Turn complex infrastructure problems into clear requirements and coherent system behavior, including APIs, lifecycle states, failure handling, upgrades, and customer controls.
  3. Partner Closely With Engineering: Work directly with engineers throughout discovery, design, implementation, and validation to resolve ambiguity and shape high-quality solutions.
  4. Establish a High Product-Quality Bar: Pressure-test designs, evaluate working software, examine edge cases and failure modes, and ensure the product is ready before broad customer release.
  5. Understand Customers and Operators: Work with customers and internal teams to identify underlying problems and translate them into broadly useful, technically sound product improvements.

Skills

Required

  • AI and Machine Learning Infrastructure: You understand how training, fine-tuning, inference, checkpointing, and GPU-intensive workloads interact with orchestration systems.
  • Strong Technical Fluency: Significant experience with Kubernetes, cloud infrastructure, distributed systems, workload orchestration, or related technologies.
  • Technical Product Management Experience: Five or more years of experience owning complex technical products or infrastructure initiatives from definition through delivery and iteration.
  • Systems Thinking and Attention to Detail: The ability to reason through APIs, lifecycle states, failure modes, upgrades, operational workflows, and dependencies across distributed systems.
  • Strong Product Judgment: The ability to balance customer value, reliability, engineering cost, operational complexity, and long-term maintainability when making product decisions.
  • Ownership and Execution: A track record of independently driving ambiguous, cross-functional initiatives to clear decisions and high-quality outcomes.
  • Clear Communication and Collaboration: The ability to work closely with engineers, document decisions precisely, and create alignment across technical and nontechnical teams.

Nice to have

  • High-Performance Computing: You have experience with HPC environments, Slurm or similar workload managers, specialized networking, or large accelerator clusters.
  • Infrastructure Operations: You have operated or supported production infrastructure and understand the impact of product decisions on reliability, upgrades, debugging, and incident response.
  • Builder Orientation: You enjoy using infrastructure products directly, examining APIs, and experimenting with systems to develop grounded product opinions.

What the JD emphasized

  • deeply technical and hands-on product role
  • complex and ambiguous infrastructure problems
  • detailed behavior
  • edge cases and failure modes
  • thoughtful product trade-offs
  • what we ship works well for customers in practice
  • raising the quality of both the product and the decisions behind it
  • technical depth
  • customer understanding
  • strong product judgment
  • disciplined execution
  • customers can depend on for their most demanding AI workloads
  • AI and Machine Learning Infrastructure
  • Significant experience with Kubernetes, cloud infrastructure, distributed systems, workload orchestration, or related technologies.
  • Five or more years of experience owning complex technical products or infrastructure initiatives from definition through delivery and iteration.
  • Systems Thinking and Attention to Detail
  • reason through APIs, lifecycle states, failure modes, upgrades, operational workflows, and dependencies across distributed systems.
  • Strong Product Judgment
  • balance customer value, reliability, engineering cost, operational complexity, and long-term maintainability
  • Ownership and Execution
  • independently driving ambiguous, cross-functional initiatives to clear decisions and high-quality outcomes.
  • Clear Communication and Collaboration

Other signals

  • AI infrastructure
  • Kubernetes
  • workload scheduling
  • customer APIs