Data & ML Pipeline Software Engineer

Applied Intuition Applied Intuition · Robotics · Sunnyvale, CA · SDS Software Engineering

Software Engineer focused on building and maintaining large-scale data processing pipelines (ETL) for ingesting and curating driving datasets, designing systems that automate data selection, labeling, training, and testing loops, and developing infrastructure to close the loop between real-world test results and new model deployments for autonomous vehicles.

What you'd actually do

  1. Build and maintain large-scale data processing pipelines (ETL) for ingesting and curating driving datasets.
  2. Design and implement systems that automate data selection, labeling, training, and testing loops.
  3. Collaborate with modeling teams to improve training efficiency and model performance across iterations.
  4. Develop the core infrastructure that closes the loop between real-world test results and new model deployments.
  5. Use your engineering expertise to help Applied Intuition’s vehicles learn from data at scale, improving safety and performance.

Skills

Required

  • Bachelor's or higher degree in Engineering such as Computer Science, Electrical Engineering, Software Engineering
  • 3–5 years of experience in software or data infrastructure engineering.
  • Expertise in building and scaling data pipelines, distributed systems, or ML infrastructure.
  • Proficiency in Python
  • strong knowledge of data frameworks (Spark, Airflow, Kafka, etc.)
  • Experience working with large-scale datasets
  • understanding data-driven development cycles
  • Familiarity with machine learning workflows or model training/deployment, especially automation of those processes.
  • Strong systems thinking
  • ability to work across multiple parts of the stack (data, infra, and ML)

Nice to have

  • Experience with automotive (AV) or robotics systems.
  • Previous work on ML platforms for large-scale products (e.g., Ads, Recommendation, or Autonomy pipelines).
  • Experience with highly automated ML training workflows.
  • Prior contributions to systems that connect data-driven model iteration loops (“data flywheel”).
  • Ability to move fast, learn quickly, and mentor others while growing with the team.

What the JD emphasized

  • automate data selection
  • training
  • testing loops
  • model iteration
  • model deployments
  • vehicles learn from data at scale

Other signals

  • building systems that connect vehicle data collection, training, and automated model improvement
  • automate data selection, curation, and model iteration
  • automate data selection, labeling, training, and testing loops
  • closes the loop between real-world test results and new model deployments
  • vehicles can self-improve with minimal human intervention