Applied Scientist - Reinforcement Learning, Omhs Scs

Amazon Amazon · Big Tech · N.reading, MA · Applied Science

This role focuses on designing, implementing, and deploying novel Reinforcement Learning (RL) agents and control policies for real-time Material Handling Equipment (MHE) control and building-wide optimization in industrial environments. It involves formulating operational problems as sequential decision-making problems, designing reward functions, building simulation environments for training and validation, and integrating RL policies into production systems.

What you'd actually do

  1. Own the research and development of reinforcement learning and sequential decision making solutions spanning deep RL, policy optimization, offline/batch RL, contextual bandits, and multi-agent RL for real-time MHE control and building-wide optimization in a production environment.
  2. Formulate fulfillment operations problems (throughput optimization, flow, merge, and congestion control) as sequential decision-making problems, and design multi-objective reward functions that balance competing operational objectives.
  3. Build and leverage high-fidelity simulation environments for safe offline training, policy validation, and sim-to-real transfer before fleet-scale deployment.
  4. Collaborate across multiple science and engineering teams to integrate RL policies into real-time production and control systems.

Skills

Required

  • PhD in computer science, machine learning, engineering, or related fields
  • 2+ years of building machine learning models or developing algorithms for business application experience
  • Demonstrated experience developing and applying reinforcement learning and sequential decision-making methods (e.g., deep RL, policy gradient / actor-critic methods, offline RL, contextual bandits, or multi-agent RL) to real-world control or optimization problems.
  • Fluency in a high-level programming language such as Python
  • Experience with popular deep learning frameworks (e.g., PyTorch, TensorFlow)
  • Experience with RL tooling or simulators (e.g., Ray/RLlib, Gymnasium, Stable-Baselines3, Isaac Gym/Omniverse, MuJoCo)
  • Ownership of end-to-end solutions in terms of research, prototyping, and experimentation.

Nice to have

  • C++
  • First-author publications at top-tier machine learning and AI venues (e.g.,NeurIPS, ICML, ICLR, AAAI, AISTATS, CoRL, or RLC/RLDM).
  • Experience applying RL in a setting analogous to ours: real-time control, robotics or material handling, industrial process or operations.
  • Experience deploying RL or ML models to production at scale and partnering with engineering teams on real-time inference and feedback loops.

What the JD emphasized

  • reinforcement learning
  • sequential decision making
  • real-time MHE control
  • building-wide optimization
  • production environment
  • multi-agent RL
  • offline/batch RL
  • contextual bandits
  • deep RL
  • policy optimization
  • multi-objective reward functions
  • high-fidelity simulation
  • sim-to-real transfer
  • fleet-scale deployment
  • real-time production
  • control systems

Other signals

  • design, implement, and deploy novel RL agents
  • real-time MHE control and building-wide optimization
  • high-fidelity simulation for safe offline training, policy validation, and sim-to-real transfer
  • integrate RL policies into real-time production and control systems