Senior Data Engineer I

Samsara Samsara · Enterprise · San Francisco, CA · Remote · Platform

Senior Data Engineer responsible for designing and maintaining data pipelines within a central data lake, utilizing SparkSQL and Pyspark. These pipelines ingest and transform data from IoT devices and software products to support statistical analysis, model training, and dashboard creation. The role involves collaborating with Data Science & Analytics and AI/ML teams.

What you'd actually do

  1. Build and maintain highly reliable computed tables, incorporating data from various sources, including unstructured data like video and audio, Samsara sensor & product data, and customer metadata.
  2. Access, manipulate, and integrate external datasets with internal data.
  3. Deliver high-quality data with strong uptime and reliability requirements, including customer-facing data sets.
  4. Collaborate closely with cross-functional teams such as Data Science & Analytics, AI/ML, and other Data Engineers to ensure high-quality data for diverse purposes from causal inference, model training, and dashboarding.
  5. Champion, role model, and embed Samsara’s cultural principles (Focus on Customer Success, Build for the Long Term, Adopt a Growth Mindset, Be Inclusive, Win as a Team) as we scale globally and across new offices.

Skills

Required

  • 4+ years experience in a data engineering-focused role
  • Designing data models at scale
  • Building ETL pipelines to handle large volumes of data
  • Spark-based data platforms
  • SQL
  • Python
  • working with REST APIs
  • software engineering fundamentals
  • reading backend development code
  • version control systems such as Git/GitHub
  • data orchestration tool (e.g Airflow, Dagster, or Prefect)

Nice to have

  • time series data
  • late-arriving data
  • Databricks
  • Delta Lakes
  • Dagster
  • public cloud (e.g AWS, GCP, Azure)
  • product’s first-party data
  • complex data, including ML outputs and/or client-side signals