Senior Data Scientist, Alexa for Shopping (rufus)

Amazon Amazon · Big Tech · Seattle, WA · Data Science

This role focuses on building and scaling a production agentic intelligence system for product analytics within Amazon Shopping CX. The core responsibilities involve multi-agent system orchestration, context management, self-improvement mechanisms, and reliable signal extraction from unstructured data. A key aspect is developing a system to automatically evaluate its own output quality and feedback loop closure.

What you'd actually do

  1. Own the multi-agent topology (Planner → Worker → Reasoner → Loop Controller) — inter-agent communication protocols, and loop termination logic
  2. Design and manage the context window strategy across agents
  3. Own all system prompts, routing prompts, and chain-of-thought scaffolding across agents
  4. Define what "better" means across dimensions (factual grounding, hypothesis novelty, evidence completeness, reasoning coherence) without ground-truth labels at scale
  5. Design how eval signal propagates back into prompt updates and model routing decisions

Skills

Required

  • 4+ years of data scientist experience
  • 5+ years of data querying languages (e.g. SQL), scripting languages (e.g. Python) or statistical/mathematical software (e.g. R, SAS, Matlab, etc.) experience
  • Experience with statistical models e.g. multinomial logistic regression
  • 5+ years of working with Data & AI related technologies, including, but not limited to, AI/ML (Artificial Intelligence/Machine Learning), GenAI (Generative AI), Analytics, Database, and/or Storage experience
  • Python proficiency — statistical modeling, data manipulation (pandas, numpy, scipy), and scripting across ML pipelines and evaluation infrastructure
  • Demonstrated experience extracting structured signal from unstructured text at scale — NLP pipelines, intent classification, entity extraction, or equivalent

Nice to have

  • Experience with multi-agent system evaluation and independently and end-to-end
  • Production RAG or retrieval system experience — embedding strategy, vector search, hybrid retrieval, similarity threshold calibration
  • AWS Bedrock or Strands SDK experience — or equivalent orchestration framework (LangGraph, CrewAI, AutoGen)
  • Graph database experience (Neptune, Neo4j) — schema design, traversal queries, knowledge graph construction
  • Experience scaling NLP inference pipelines — model sizing decisions, batching strategy, SageMaker or equivalent endpoint optimization
  • Business intelligence or analytics domain background — metric definitions, dimensional modeling, causal inference
  • Track record of publishing at peer-reviewed venues or presenting at industry conferences

What the JD emphasized

  • operating independently on problems that are not well-defined or structured
  • identifying and framing research challenges across broad problem areas
  • delivering end-to-end solutions that have significant impact on the product
  • evaluates its own output quality
  • identifies where it fails
  • closes that feedback loop automatically

Other signals

  • agentic intelligence system
  • multi agent system orchestration
  • self-improving agent layer
  • reliable signal extraction
  • evaluates its own output quality
  • identifies where it fails
  • closes that feedback loop automatically