Sr. Research Engineer

Adobe Adobe · Enterprise · San Francisco, CA +1

Research Engineer role focused on owning the end-to-end training data pipelines for generative audio models, including ingestion, filtering, preprocessing, curation, acquisition, and annotation. The role involves running inference, building data exploration tools, training/finetuning models, and evaluating their behavior to drive data experiments. It requires strong audio domain knowledge, generative ML research experience, and data acquisition/licensing skills.

What you'd actually do

  1. Build, deploy, and monitor large-scale data pipelines for ingesting, filtering and preprocessing audio and video at scale, with reproducibility and versioning.
  2. Run in-house models for inference over millions of audio/video files, and build data exploration tools to support this.
  3. Curate and select training data using both quantitative metrics and qualitative judgment. Maintain an ongoing understanding of the corpus: distribution, balance, diversity, coverage, and redundancy.
  4. Train or finetune models on pipeline outputs, evaluate their behavior, and use those findings to drive new data experiments and ablations. Scale training and optimize throughput.
  5. Drive data licensing and acquisition: identify gaps and opportunities, define briefs, and work with vendors and producers to license and curate new datasets.

Skills

Required

  • Pipeline-building experience at scale for audio and/or video data with reproducibility and versioning
  • Deep audio domain knowledge: common datasets, quality metrics and their failure modes, codecs, normalization, and filtering. Strong "ears" and critical listening ability. Video knowledge a plus.
  • Generative-ML research experience to make data decisions, design and run training experiments independently
  • Strong command of evaluation in audio and video modelling.
  • Data acquisition and licensing experience, including writing briefs and communicating across stakeholders.
  • Excellent communication skills

Nice to have

  • Video knowledge

What the JD emphasized

  • own the training data
  • single owner of _what's in the data, why, and how we know it's good.
  • pipeline-building experience at scale
  • Deep audio domain knowledge
  • Generative-ML research experience
  • Strong command of evaluation

Other signals

  • audio GenAI
  • generative audio models
  • training data
  • model performance
  • productizing research