Software Engineer - X Data

xAI xAI · AI Frontier · Palo Alto, CA · Engineering

xAI is seeking a Data Engineer to build and operate a distributed data platform that processes billions of events daily through real-time and batch pipelines. The role involves creating shared datasets, internal data products, and automation tooling for data workflows. Responsibilities include owning data correctness end-to-end, utilizing various query engines and frameworks, partnering with product and business teams, and iterating quickly on feedback to deliver reliable solutions.

What you'd actually do

  1. Design, build, and operate production-grade realtime and batch pipelines that ingest, process, validate, and deliver data powering user-behavior insights and product decisions.
  2. Create shared datasets, fact tables, and internal data products that let other teams analyze, debug, and improve product performance.
  3. Prototype and build tooling that automates and accelerates internal data workflows — backfills, dashboards, report generation, and self-serve access to data.
  4. Own data correctness end to end: validate with output invariants, denominator reconciliation, and independent recomputation, and lead root-cause investigations when key metrics move unexpectedly.
  5. Move fluidly across query engines and frameworks (e.g., BigQuery, Trino, Clickhouse for analytics; Flink, Kafka, Spark/Scalding for streaming and batch), choosing the right tool and adapting quickly to new infrastructure and environments.

Skills

Required

  • Python
  • Rust
  • Scala
  • Go
  • Java
  • data pipeline toolings
  • distributed systems
  • Spark
  • Kafka
  • Flink
  • SQL
  • RMDBs
  • NoSQL
  • large scale problems
  • incremental quality work
  • building brand new systems
  • interpreting product requirements into engineering implementation plans
  • communicating with different groups

What the JD emphasized

  • own problems end to end
  • own data correctness end to end
  • shipping the smallest useful increment