Sr. Engineering Manager, Inference

Weights & Biases Weights & Biases · Data AI · Bellevue, WA +5 · Remote · Technology

Senior Engineering Manager for AI/ML Platform team at CoreWeave, focusing on productizing and operating their inference offering. The role involves leading a team to ensure service reliability, observability, operational excellence, and developer experience for AI workloads. Responsibilities include roadmap execution, engineering processes, cross-functional partnerships, and building a culture of technical excellence.

What you'd actually do

  1. Lead and grow the engineering team responsible for evolving and operating the W&B Inference product, focusing on service reliability, orchestration, operational maturity, and developer experience.
  2. Drive execution on roadmap initiatives in close partnership with Product, ensuring that platform capabilities are delivered predictably, robustly, and with measurable customer impact.
  3. Own the engineering processes for the inference productization layer, including incident response, operational readiness, release management, observability, and engineering quality.
  4. Partner with CoreWeave’s infrastructure teams to integrate capabilities from the underlying inference platform into a cohesive, user-facing product.
  5. Guide the design and delivery of application-layer enhancements such as tracing and tool call handling.

Skills

Required

  • 7+ years of experience in software engineering
  • 3+ years managing or leading engineering teams responsible for distributed systems, developer platforms, or large-scale services
  • Fluent in concepts related to distributed compute services, including observability, autoscaling patterns, reliability engineering, API design, and service operations
  • Skilled at leading engineering teams through ambiguous, multi-stakeholder projects with strong communication and alignment
  • Deep empathy for ML practitioners and platform developers

Nice to have

  • Background in high-scale systems, real-time APIs, or cloud infrastructure, ideally with exposure to model-serving or inference-adjacent domains
  • Background in related platform domains such as IAM, billing/metering, observability systems, or deployment tooling
  • Curious about AI and MLOps tooling
  • Expert in building inference systems that scale for production workloads

What the JD emphasized

  • productizing and operating our inference offering
  • service reliability
  • observability
  • operational excellence
  • developer-facing tooling
  • inference experience
  • practitioners deploying and scaling real-world AI workloads
  • roadmap initiatives
  • platform capabilities
  • customer impact
  • engineering processes
  • incident response
  • operational readiness
  • release management
  • engineering quality
  • underlying inference platform
  • application-layer enhancements
  • tracing
  • tool call handling
  • reliability
  • usability
  • compliance
  • performance
  • cost
  • operational load
  • architectural constraints
  • complex, cross-functional environments
  • communication
  • aligned priorities
  • smooth execution
  • multiple teams
  • culture of ownership
  • technical excellence
  • continuous improvement
  • distributed systems
  • developer platforms
  • large-scale services
  • distributed compute services
  • autoscaling patterns
  • reliability engineering
  • API design
  • service operations
  • competing priorities
  • principled trade-offs
  • latency
  • development velocity
  • ambiguous, multi-stakeholder projects
  • ML practitioners
  • platform developers
  • reduce friction
  • elevate the developer experience
  • high-scale systems
  • real-time APIs
  • cloud infrastructure
  • model-serving
  • inference-adjacent domains
  • IAM
  • billing/metering
  • observability systems
  • deployment tooling
  • build frictionless products for developers
  • AI and MLOps tooling
  • building inference systems that scale for production workloads

Other signals

  • leading the group responsible for productizing and operating our inference offering
  • ensure the service becomes a polished, reliable, developer-friendly product
  • guide engineers working on areas such as service reliability, observability, operational excellence, packaging, developer-facing tooling, and application-layer enhancements
  • partner closely with Product, CoreWeave engineering teams, Design, Support, and GTM stakeholders to ensure the inference experience meets the needs of practitioners deploying and scaling real-world AI workloads
  • drive execution on roadmap initiatives in close partnership with Product, ensuring that platform capabilities are delivered predictably, robustly, and with measurable customer impact
  • own the engineering processes for the inference productization layer, including incident response, operational readiness, release management, observability, and engineering quality
  • guide the design and delivery of application-layer enhancements such as tracing and tool call handling
  • ensure the inference offering meets high standards of reliability, usability, compliance, and performance
  • build a culture of ownership, technical excellence, and continuous improvement within the engineering team
  • 7+ years of experience in software engineering, including 3+ years managing or leading engineering teams responsible for distributed systems, developer platforms, or large-scale services
  • fluent in concepts related to distributed compute services, including observability, autoscaling patterns, reliability engineering, API design, and service operations
  • skilled at leading engineering teams through ambiguous, multi-stakeholder projects with strong communication and alignment
  • deep empathy for ML practitioners and platform developers, with a drive to improve reliability, reduce friction, and elevate the developer experience
  • background in high-scale systems, real-time APIs, or cloud infrastructure, ideally with exposure to model-serving or inference-adjacent domains
  • expert in building inference systems that scale for production workloads