Applied Scientist, Jp Ops Star

Amazon Amazon · Big Tech · 13, Japan +1 · Software Development

Applied Scientist role focused on building and deploying LLM-based agent systems and graph neural networks for Amazon's transportation operations. The role involves end-to-end ownership from formulation to production, with a focus on GPU-optimized training and inference, and rigorous evaluation frameworks.

What you'd actually do

  1. Formulate and solve modeling problems for LLM-powered agent systems (MCP servers, text-to-SQL, RAG), designing validation methods and confidence scoring for reliable outputs under operational constraints.
  2. Design evaluation frameworks for AI agents: define metrics for reasoning quality and task completion, build backtesting pipelines, and develop automated regression detection.
  3. Research and implement graph neural network architectures to learn structural representations of logistics networks; develop training methodology and evaluation protocols for production deployment.
  4. Develop GPU-optimized pipelines for model training and inference, with rigorous benchmarking of latency, throughput, and cost trade-offs.
  5. Own the scientific lifecycle end-to-end: problem formulation, experimental design, deployment, and continuous measurement against real operational outcomes.

Skills

Required

  • building models for business application
  • large scale machine learning systems
  • training and deploying machine learning systems
  • Java, C++, Python or related language

Nice to have

  • patents or publications at top-tier peer-reviewed conferences or journals
  • developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware
  • computer architecture

What the JD emphasized

  • LLM-based agent systems
  • GPU-optimized model serving and training
  • Graph representation learning
  • production deployment
  • scientific rigor
  • well-designed experiments
  • proper baselines
  • quantified uncertainty
  • engineering judgment
  • publish and share their work
  • treat production performance as ground truth

Other signals

  • LLM-based agent systems
  • GPU-optimized model serving and training
  • Graph representation learning
  • production deployment