Senior Manager, Production Engineering

Weights & Biases Weights & Biases · Data AI · Bellevue, WA +4 · Technology

Senior Manager of Production Engineering to lead and expand the SRE team and practices, focusing on the reliability, performance, and operational efficiency of a large-scale, distributed cloud infrastructure, with experience in AI infrastructure environments and GPU-accelerated workloads.

What you'd actually do

  1. Contribute and execute the SRE vision, strategy, and roadmap for a large-scale, distributed cloud infrastructure.
  2. Lead and mentor a high-performing team of SREs, promoting a culture of ownership, collaboration, and continuous learning.
  3. Champion automation-first practices, leveraging AI, tools like Terraform, Kubernetes, and Infrastructure-as-Code to minimize toil and manual interventions.
  4. Establish and evolve Operational Excellence best practices ensuring the platform is proactive, and propagates organizational learning.
  5. Drive initiatives for incident management, postmortem culture, root cause analysis, and system hardening.

Skills

Required

  • Leadership
  • SRE
  • Incident Management
  • Distributed Systems
  • Networking
  • Storage Architecture
  • Cross-functional collaboration

Nice to have

  • Platform tooling
  • Internal developer portals
  • GPU-accelerated workloads
  • AI infrastructure environments
  • Compliance
  • Reliability risk modeling
  • Bare metal infrastructure
  • DPUs
  • Service mesh architectures
  • Multi-tenant security models

What the JD emphasized

  • 10+ years in a leadership or senior management role
  • Experience hiring, developing, and managing geographically distributed 24x7 engineering teams.
  • Experience in designing and implementing incident management processes including on-call rotations, escalation paths, postmortems, and SLO/SLA framework.