Sr Applied Scientist, Core Shopping Data Science

Amazon Amazon · Big Tech · Seattle, WA · Applied Science

This role focuses on bridging the gap between LLM-driven quality judgments and production-level ranking systems in a large-scale consumer shopping environment. The scientist will distill LLM quality judgments into compact, efficient models suitable for online serving under strict latency and cost constraints. They will also model customer responses to perceived defects and work with ranking teams to integrate these signals into online objectives, ultimately driving improvements in customer experience visible to nearly every Amazon customer.

What you'd actually do

  1. You will model how customers respond to the defects they encounter. Using newly instrumented logging data, you will build representations of customer sessions from page sequences, intent signals, and quality exposures, and quantify what changes downstream when a customer meets an irrelevant recommendation, a set of near-duplicate widgets, or a confusing label early in a journey. You will identify the contextual factors that mediate that impact, including session intent, category, device, and prior interactions, and turn the results into a ranking of which defects are worth trading engagement to prevent, on which surfaces, for which customers.
  2. You will distill LLM quality judgments into models compact enough to serve online. That means training compact models against LLM-generated labels, characterizing where the student diverges from its teacher and on which segments, and holding accuracy under the latency and cost budgets of Search, Homepage, and Detail Page ranking. Where a distilled model cannot meet that bar, you will say so early and propose what would.
  3. You will work with Search, Homepage, and Detail Page ranking teams to get those signals into online objectives. You will analyze which of those systems offers the most leverage, recommend where to invest first, and design quality-aware objective formulations that trade impression quality against engagement deliberately, replacing the current pattern of suspending a strategy after a problem surfaces.
  4. You will prove all of it in controlled experiments. You will design the experiment, choose the right success metrics metrics and make the call on what ships.

Skills

Required

  • building machine learning models for business application
  • PhD, or Master's degree and 6+ years of applied research experience
  • programming in Java, C++, Python or related language
  • neural deep learning methods
  • machine learning

Nice to have

  • modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.
  • large scale distributed systems such as Hadoop, Spark etc.

What the JD emphasized

  • distill LLM quality judgments into models compact enough to serve online
  • under latency and cost budgets
  • model how customers respond to the defects they encounter
  • quantify what changes downstream
  • turn the results into a ranking of which defects are worth trading engagement to prevent
  • training compact models against LLM-generated labels
  • holding accuracy under the latency and cost budgets of Search, Homepage, and Detail Page ranking
  • prove all of it in controlled experiments

Other signals

  • LLMs changed that.
  • We can now take an ambiguous statement about customer perception, make it a judgment that holds up consistently at scale, and validate it against what shoppers actually do next.
  • What we cannot do is call a large model in a ranking request path at large scale.
  • Every shopping session passes through those systems under latency and cost budgets that leave no room for one.
  • So we can describe quality far better than we can optimize for it, and closing that gap is what this role exists to do.