Every day, Amazon Japan delivers millions of packages to customers' doors. Behind every routing decision, capacity plan, and network design is a modeling problem — and the science behind how those models learn, generalize, and improve in production is where you come in. JP OPS STAR Foundation is the applied science and engineering team that builds the decision systems powering Amazon Japan's transportation operations. We sit at the intersection of generative AI, large-scale optimization, and graph-based learning — developing models that turn complex operational structure into actionable intelligence. As an applied scientist, you will formulate and solve research problems that directly shape operational outcomes. Your work will span:
- LLM-based agent systems — designing agentic architectures with structured evaluation frameworks that measure reasoning quality, tool use accuracy, and task completion under real operational constraints
- GPU-optimized model serving and training — developing pipelines for large model inference and fine-tuning, with rigorous benchmarking of latency, throughput, and cost trade-offs
- Graph representation learning — building graph neural network models that capture logistics network topology and learn node/edge representations for downstream prediction and optimization tasks You will own problems end-to-end: from formulation and experimentation through deployment and continuous measurement in production. We expect scientific rigor — well-designed experiments, proper baselines, quantified uncertainty — and the engineering judgment to make your methods work reliably at scale. This is a high-autonomy, high-impact role. You will collaborate with data engineers, BIEs, and operations leaders, and you will have the freedom to define your research agenda within our problem space. We value scientists who publish and share their work, and who treat production performance as ground truth for their ideas.
At Amazon, you'll work alongside the latest AI and GenAI tools that are increasingly woven into how teams operate: from AI-powered capabilities that accelerate decision-making, to Generative AI that helps you focus on work that truly matters. You'll have opportunities and resources to develop AI fluency at your own pace, with continuous learning built into the culture.
Key job responsibilities
- Formulate and solve modeling problems for LLM-powered agent systems (MCP servers, text-to-SQL, RAG), designing validation methods and confidence scoring for reliable outputs under operational constraints.
- Design evaluation frameworks for AI agents: define metrics for reasoning quality and task completion, build backtesting pipelines, and develop automated regression detection.
- Research and implement graph neural network architectures to learn structural representations of logistics networks; develop training methodology and evaluation protocols for production deployment.
- Develop GPU-optimized pipelines for model training and inference, with rigorous benchmarking of latency, throughput, and cost trade-offs.
- Own the scientific lifecycle end-to-end: problem formulation, experimental design, deployment, and continuous measurement against real operational outcomes.
- Collaborate with data engineers and operations stakeholders to validate model outputs and quantify business impact.
A day in the life
- Diagnose a confidence-score drift in the agent eval dashboard, trace the root cause to a schema change, and redesign the validation logic to be schema-invariant.
- Experiment with a graph attention architecture on the logistics network; benchmark against the baseline embedding on a downstream prediction task.
- Profile a GPU training job to unblock a scaling experiment, identify a memory bottleneck, and restructure the data loader to halve iteration time.
- Present model accuracy trends to operations leaders and recommend threshold adjustments based on the data.
Basic Qualifications
- 3+ years of building models for business application experience
- PhD, or Master's degree and 4+ years of CS, CE, ML or related field experience
- Experience programming in Java, C++, Python or related language
- Experience with large scale machine learning systems such as profiling and debugging and understanding of system performance and scalability
- Experience with training and deploying machine learning systems to solve large-scale optimizations
Preferred Qualifications
- Experience in patents or publications at top-tier peer-reviewed conferences or journals
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience in computer architecture
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.