Software Engineer, Data Operations

Superhuman Superhuman · Consumer · Hub - Berlin · Engineering, Product, Design, and Marketing

Software Engineer on the Data Operations team responsible for shaping the architecture and technical strategy of the data platform, ensuring scalability, security, and efficiency. Designs and leads implementation of robust systems for large data volumes, supporting product features and data-driven decision-making. Work spans real-time & ETL data pipelines, cloud infrastructure, data lakes, and back-end services. Collaborates with cross-functional teams including ML teams and mentors engineers.

What you'd actually do

  1. Architect and lead the development of large-scale systems for data pipelines, data lakes that handle billions of daily events.
  2. Design and implement solutions that ensure data is available, secure, and scalable across the platform, enabling real-time and batch processing, including ML research use cases.
  3. Make high-level architectural decisions about system design, technology choices, and platform evolution, ensuring scalability and long-term sustainability.
  4. Collaborate with key stakeholders, such as product teams, data engineers, back-end developers, and ML engineers, to build tools & frameworks that power analytics, product features, and data-driven workflows.
  5. Ensure easy of use and cost effectiveness of company wide data infrastructure

Skills

Required

  • SQL
  • Spark
  • Kafka
  • Terraform
  • Python
  • Scala
  • Java
  • Delta Lake
  • Snowflake
  • BigQuery
  • Redshift
  • data governance
  • GDPR/CCPA

Nice to have

  • large-scale distributed computing systems
  • data infrastructure provisioning
  • managing live production environments
  • high-load systems
  • data-intensive workflows
  • strategic thinker
  • build internally or leverage third-party solutions
  • excellent communicator and collaborator
  • translates business needs into robust data solutions
  • alignment on technical goals
  • work independently with minimal to zero guidance
  • proactively manages tasks and priorities
  • analyzes and executes work efficiently
  • collaborates effectively with cross-functional teams
  • thrives in fast-paced, results-driven environments

What the JD emphasized

  • large-scale systems
  • billions of daily events
  • ML research use cases
  • scalable
  • secure
  • real-time and batch processing
  • data governance
  • privacy regulations (GDPR/CCPA)
  • cost effectiveness