Reliability Lead, Common Services

Weights & Biases Weights & Biases · Data AI · Bellevue, WA +2 · Technology

CoreWeave is seeking a Reliability Lead for their Common Services organization. This role will establish and lead the Reliability Engineering and production operations practice for shared platforms, APIs, and foundational services that power their AI cloud products. Responsibilities include defining reliability strategy, processes, and standards, managing incidents, driving observability, designing for reliability, and automating operational workflows. The role requires strong experience in SRE, distributed systems, Linux, observability stacks, incident response, and infrastructure-as-code.

What you'd actually do

  1. Establish and lead the SRE / production engineering practice for the Common Services organization, including standards for reliability, incident management, and on-call, in partnership with the central Product Engineering organization.
  2. Develop an Operational Excellence strategy that focuses on not only improving system performance but also monitoring and reducing operational toil
  3. Partner with engineering and product teams to define SLOs, SLIs, and error budgets for critical Common Services, and ensure these become part of how teams plan and make tradeoffs.
  4. Own and improve the incident management lifecycle for Common Services, including on-call rotations, escalation paths, incident tooling, post-incident reviews, and follow-through on corrective actions.
  5. Drive the observability strategy (metrics, logs, traces, dashboards, alerts) for Common Services, ensuring we have actionable visibility into the health, performance, and capacity of key systems.

Skills

Required

  • Site Reliability Engineering
  • Production Engineering
  • distributed systems
  • cloud/platform services
  • technical leadership
  • Linux-based production environments
  • containers
  • Kubernetes
  • observability stacks
  • metrics
  • logging
  • tracing
  • alerting systems
  • SLIs/SLOs
  • incident response
  • capacity planning
  • redundancy
  • failover
  • backoff
  • circuit breaking
  • graceful degradation
  • infrastructure-as-code
  • automation tooling
  • Terraform
  • Ansible
  • Helm
  • CI/CD pipelines
  • cross-functional communication

Nice to have

  • GPU workloads
  • high-performance computing
  • latency/throughput-sensitive systems
  • multi-tenant environments
  • multi-region environments
  • highly regulated environments
  • service ownership models
  • mentoring senior engineers
  • building high-performing teams

What the JD emphasized

  • reliability
  • operational excellence
  • incident management
  • observability
  • reliability
  • operational risk
  • continuous improvement
  • learning from incidents
  • humane on-call
  • reliability
  • operational improvements
  • observability stacks
  • SLIs/SLOs
  • alert strategies
  • incident response
  • post-incident reviews
  • design for reliability
  • automation tooling
  • reliability considerations