Data Scientist Ii, ML Infrastructure

Pinterest Pinterest · Consumer · San Francisco, CA · Data Engineering

This role focuses on advancing the science and systems behind ML measurement, feature understanding, and causal inference at scale. The work spans areas such as production feature importance platforms, observational causal estimation in Pytorch, large-scale proxy metric development, and data-driven approaches to ML infrastructure efficiency. The individual contributor will drive foundational innovations, own the end-to-end design of production ML systems, establish rigorous methodological standards, and partner cross-functionally to turn successful research into durable platform capabilities that raise the ceiling for the entire ML organization.

What you'd actually do

  1. Translate research-grade DS workflows (e.g., proxy metrics, staleness models) into production ML pipelines using Airflow, WandB & Ray while establishing reusable patterns for other teams.
  2. Apply and productionize causal inference methods using the production ML stack (propensity scoring, IPW, TMLE) to address high-stakes measurement questions beyond experimental capabilities. Build self-serve tooling to empower non-experts to derive rigorous causal insights at scale.
  3. Partner with ML engineers and product teams to identify opportunities for improved tooling, metrics, and measurement methods, unlocking step-change improvements in model quality and business outcomes.
  4. Leverage Pinterest's rich metadata and engagement signals to build data-driven frameworks, from feature importance to content deindexing, that improve platform efficiency and speed.
  5. Design and build centralized ML platform tooling to improve feature and model creation, evaluation, and trust, including production systems that operate daily at scale across all models.

Skills

Required

  • Python
  • PyTorch
  • distributed compute (Spark, Ray)
  • workflow management tools (Airflow, Prefect, Jenkins, or similar)
  • software development best practices
  • ML theory knowledge

Nice to have

  • Ray specifically is a strong plus

What the JD emphasized

  • production ML pipelines
  • causal inference
  • ML infrastructure efficiency
  • production ML systems
  • ML platform tooling

Other signals

  • ML infrastructure efficiency
  • production ML pipelines
  • causal inference at scale
  • ML platform tooling