Senior Software Engineer, Data Lake

Robinhood Robinhood · Fintech · Bellevue, WA +3 · ENG Infrastructure

Robinhood is seeking a Senior Software Engineer for their Data Lake Ingestion team. This role focuses on building and scaling data ingestion systems, including CDC, snapshot, and streaming pipelines using technologies like Kafka and Flink. The engineer will improve pipeline reliability, data quality, and implement security measures for sensitive data. The role involves designing next-generation ingestion systems, migrating legacy pipelines, and mentoring junior engineers.

What you'd actually do

  1. Design and scale CDC, snapshot, and streaming ingestion systems that transfer operational data into the data lake.
  2. Improve pipeline reliability, scale, data quality, and monitoring to maintain ingestion health and reduce replication lag.
  3. Develop self-service onboarding flows and tooling to help partner teams register databases and schemas quickly and securely.
  4. Collaborate with platform teams and data consumers to select and implement appropriate ingestion patterns for their latency and compliance needs.
  5. Implement access controls, PII-handling workflows, and retention pipelines to protect sensitive data across its lifecycle.

Skills

Required

  • Java, Scala, or Python
  • relational databases
  • change data capture (CDC) mechanisms
  • Kafka
  • Flink
  • AWS S3
  • Kubernetes
  • schema governance

Nice to have

  • building backend platforms
  • distributed data systems
  • high-performance code
  • large-scale streaming systems
  • Spark

What the JD emphasized

  • building backend platforms or distributed data systems
  • large-scale streaming systems
  • relational databases (such as Postgres) and change data capture (CDC) mechanisms
  • AWS S3
  • schema governance