Staff Software Engineer, Storage Platform

Robinhood Robinhood · Fintech · Bellevue, WA +2 · ENG Infrastructure

Robinhood is seeking a Staff Software Engineer for their Storage Platform team. This role involves designing, implementing, and operating distributed storage systems, including sharding, query routing, and database proxy services. The engineer will focus on improving availability, failover, disaster recovery, and control plane automation for thousands of databases and caching clusters. Key responsibilities include optimizing performance, throughput, and cost efficiency, and partnering with other teams to standardize storage access while meeting regulatory and security requirements. The role requires deep expertise in PostgreSQL, distributed systems, Go/Rust, Kubernetes, AWS managed services, and observability tooling.

What you'd actually do

  1. Design and implement distributed storage systems, including sharding strategies, query routing layers, and database proxy services
  2. Improve availability, failover, and disaster recovery capabilities across multi-region deployments
  3. Build and enhance control plane automation that manages thousands of databases and caching clusters at scale
  4. Optimize database performance and throughput by analyzing query patterns, improving connection pooling, and implementing observability standards
  5. Partner with product and infrastructure teams to standardize how services connect to Postgres, NoSQL, and caching systems while meeting regulatory and security requirements

Skills

Required

  • Deep expertise in PostgreSQL (Aurora PostgreSQL preferred) and strong understanding of relational database internals
  • Experience building or operating distributed systems, including horizontal sharding, replication, and distributed transactions (e.g., two-phase commit)
  • Proficiency in Go and/or Rust, with experience building backend or infrastructure services
  • Hands-on experience with Kubernetes, AWS managed services (RDS, DynamoDB), and observability tooling (metrics, logging, tracing)
  • Strong understanding of system reliability, performance tuning, and cost optimization in high-scale production environments

What the JD emphasized

  • Availability is our highest priority
  • meet strict uptime targets
  • no downtime during market hours
  • regulatory and security requirements