Staff Software Engineer, Observability

Robinhood Robinhood · Fintech · Toronto, ON · ENG Infrastructure

Staff Software Engineer on the Observability team at Robinhood, responsible for leading the technical roadmap and architecture of the full-stack observability platform, including metrics, logs, traces, and alerting. The role involves owning the control plane, managing telemetry ingestion costs, building scalable solutions, and integrating AI-driven approaches to ensure high availability and reliability of Robinhood's infrastructure.

What you'd actually do

  1. Define and execute the full-stack observability roadmap, establishing a clear technical vision for metrics, logs, traces, and alerting infrastructure across Robinhood's engineering organization.
  2. Own and evolve the observability control plane, including telemetry ingestion pipelines, cost management strategies, and the tooling that ensures observability components remain highly available.
  3. Lead and collaborate with a team of six engineers to build scalable, self-service observability solutions that enable product and infrastructure teams to move faster with greater confidence.
  4. Establish and maintain SLOs for observability systems that meet or exceed Robinhood's 99.9% uptime target, ensuring the observability platform is as reliable as the services it monitors.
  5. Partner with the Robinhood Command Center and engineering teams across the organization to align on dependency mapping, incident response workflows, and observability standards.

Skills

Required

  • 8+ years of software engineering experience
  • owning and delivering large-scale observability or infrastructure platform initiatives
  • Deep expertise in Kubernetes
  • public cloud environments (AWS preferred)
  • architect and operate distributed systems at scale
  • Strong coding proficiency in one or more languages (Go, Python, or similar)
  • integrating observability agents, libraries, and instrumentation directly into production codebases
  • owning an observability control plane or telemetry pipeline
  • ingestion cost management
  • cardinality control
  • signal routing

Nice to have

  • Experience with the Vector data pipeline
  • high-throughput log/metrics routing tools
  • Prometheus
  • Grafana
  • Honeycomb
  • Humio
  • Sentry

What the JD emphasized

  • observability platform
  • telemetry ingestion
  • cost management
  • high-traffic production environment
  • 99.9% uptime