Principal Scientist - Data Pipeline Engineer

Adobe Adobe · Enterprise · San Jose, CA +2

This role focuses on architecting and scaling multimodal data processing pipelines and infrastructure for Adobe Firefly's foundation models. The primary responsibility is to build distributed, GPU-accelerated systems that transform raw assets into training-ready data, directly impacting model learning speed and quality. The role also involves scaling inference throughput, optimizing data infrastructure, and driving data curation for model training, with a strong emphasis on large-scale data engineering and ML infrastructure.

What you'd actually do

  1. Architect and optimize large-scale distributed pipelines that process billions of images, video, and audio assets through ML workflows into training-ready data
  2. Scale up inference throughput across the pipeline (batching, parallelism, hardware utilization) to turn raw collected data into training data faster and more cheaply
  3. Identify and eliminate bottlenecks across ingestion, processing, and delivery, from storage and I/O to compute scheduling
  4. Design systems that reliably store, index, and serve billions of data points, each requiring substantial processing spanning large-scale databases, distributed storage, and high-throughput compute
  5. Apply deep expertise in distributed systems and frameworks such as Ray (or equivalent) to orchestrate large-scale, GPU/CPU-heavy data workloads

Skills

Required

  • 10+ years of experience in data engineering, ML infrastructure, or distributed systems
  • Hands-on expertise in distributed systems and frameworks such as Ray, Spark, or equivalent
  • Proficiency in Python
  • Strong experience in a systems-level language (C++, Rust, Go, or Java)
  • Strong debugging skills across distributed and ML-centric runtime environments
  • Deep knowledge of databases and storage systems at scale
  • Strong ML background, particularly expertise in optimizing GPU inference pipelines for VLMs, LLMs, or other large models
  • Experience with data curation for model training

Nice to have

  • architect and scale multimodal data processing pipelines
  • behind Adobe Firefly’s multimodal foundation models
  • turning billions of raw assets into training-ready data at scale
  • directly determine how fast and how well Adobe models can learn
  • broad technical influence across data, infrastructure, and modeling teams
  • large-scale distributed pipelines
  • billions of images, video, and audio assets
  • training-ready data at scale
  • Scale up inference throughput
  • billions of data points
  • large-scale databases
  • distributed storage
  • high-throughput compute
  • large-scale, GPU/CPU-heavy data workloads
  • scale alongside data and model growth
  • data curation for model training
  • optimizing GPU inference pipelines for VLMs, LLMs, or other large models
  • data curation for training generative or multimodal models
  • Comfort operating across the full stack, from low-level systems and GPU optimization to higher-level data strategy and curation decisions
  • Ability to communicate clearly and partner effectively across data, infrastructure, and modeling teams
  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Machine Learning, or a related field

What the JD emphasized

  • architect and scale the multimodal data processing pipelines and infrastructure
  • behind Adobe Firefly’s multimodal foundation models
  • turning billions of raw assets into training-ready data at scale
  • directly determine how fast and how well Adobe models can learn
  • broad technical influence across data, infrastructure, and modeling teams
  • large-scale distributed pipelines
  • billions of images, video, and audio assets
  • training-ready data at scale
  • Scale up inference throughput
  • billions of data points
  • large-scale databases
  • distributed storage
  • high-throughput compute
  • large-scale, GPU/CPU-heavy data workloads
  • scale alongside data and model growth
  • data curation for model training
  • optimizing GPU inference pipelines for VLMs, LLMs, or other large models
  • data curation for training generative or multimodal models
  • 10+ years of experience in data engineering, ML infrastructure, or distributed systems, including work at large scale (billions of records or assets)
  • hands-on expertise in distributed systems and frameworks such as Ray, Spark, or equivalent large-scale data processing frameworks
  • strong experience in a systems-level language (C++, Rust, Go, or Java) with strong debugging skills across distributed and ML-centric runtime environments
  • Deep knowledge of databases and storage systems at scale such as data lakes, indexing, and retrieval across billions of data points
  • Strong ML background, particularly expertise in optimizing GPU inference pipelines for VLMs, LLMs, or other large models (batching, quantization, serving, throughput/latency tradeoffs)
  • Experience with data curation for model training: understanding what makes data valuable for training generative or multimodal models, not just how to move it efficiently

Other signals

  • architect and scale multimodal data processing pipelines
  • behind Adobe Firefly’s multimodal foundation models
  • turning billions of raw assets into training-ready data at scale
  • directly determine how fast and how well Adobe models can learn