Senior Software Engineer, Robinhood Command Center

Robinhood Robinhood · Fintech · Menlo Park, CA +1 · ENG Technical Assurance

Senior Software Engineer for Robinhood's Command Center (RCC), a new reliability team focused on detecting, coordinating, and mitigating production incidents. The role involves driving reliability and observability strategy, leading incident mitigation, developing incident management processes and tooling, owning incident discovery and response tooling, driving post-incident learning, designing failure mitigation strategies, and improving monitoring/alerting. Requires strong experience in operating production systems, reliability engineering, incident leadership, and observability frameworks.

What you'd actually do

  1. Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure
  2. Partner closely across many different types of engineers to raise the bar for operational excellence and incident response
  3. Lead incident mitigation efforts by coordinating service owners, facilitating time-sensitive decisions like rollbacks, traffic shifts, and maintaining a clear source of truth during active incidents
  4. Develop and maintain incident management processes and procedures to ensure timely resolution and minimize customer impact
  5. Own incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys (CUJs), availability, and business-impact metrics

Skills

Required

  • 5+ years of software engineering experience
  • significant experience operating production systems
  • 2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations
  • Hands-on experience serving in incident leadership roles (e.g., IMOC, incident commander, primary oncall)
  • Strong communication and cross-functional collaboration skills, especially during high-severity incidents
  • Deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design
  • Experience with multi-region or multi-cluster architectures, capacity planning, and failover strategies
  • Familiarity with modern observability stacks (e.g., OpenTelemetry, Prometheus, Grafana)
  • Demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact

What the JD emphasized

  • long-term reliability and observability strategy
  • incident leadership
  • operational excellence
  • incident response
  • incident discovery
  • incident response tooling and processes
  • post-incident governance and learning
  • failure mitigation strategies
  • monitoring, alerting, and observability
  • observability to critical user journeys
  • service quality and reliability
  • operating production systems
  • reliability engineering
  • production operations
  • incident leadership roles
  • high-severity incidents
  • systems reliability
  • observability frameworks
  • fault-tolerant architecture design
  • multi-region or multi-cluster architectures
  • capacity planning
  • failover strategies
  • modern observability stacks
  • measurable improvements in MTTD, MTTR, availability, or customer impact