Senior Applied Scientist, Network Fabric Engineering

Amazon Amazon · Big Tech · Seattle, WA · Applied Science

This role focuses on developing novel optimization algorithms and ML models for network traffic engineering and capacity planning at planetary scale within AWS Inter Data Center Networking. It involves formulating problems as mathematical optimization models, designing ML models for prediction and anomaly detection, building capacity planning frameworks, and creating simulation/evaluation frameworks. The role requires collaboration with engineers and operations teams to productionize algorithms and publish research. The primary output is the ML models and optimization algorithms, with a secondary focus on data pipelines for these models.

What you'd actually do

  1. Formulate traffic engineering problems as mathematical optimization models and develop scalable solvers that operate on graphs with hundreds of nodes and O(10^4) edges
  2. Design ML models for traffic demand prediction, anomaly detection, and workload classification
  3. Develop capacity planning frameworks that co-optimize cost, reliability, and performance over multi-year horizons
  4. Build simulation and evaluation frameworks to validate TE algorithms against realistic failure scenarios and traffic patterns
  5. Publish research at top venues (SIGCOMM, NSDI, INFOCOM, NeurIPS, ICML) and contribute to the scientific community

Skills

Required

  • building machine learning models for business application experience
  • PhD, or Master's degree and 6+ years of applied research experience
  • Experience programming in Java, C++, Python or related language
  • Experience with neural deep learning methods and machine learning

Nice to have

  • modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.
  • large scale distributed systems such as Hadoop, Spark etc.

What the JD emphasized

  • building machine learning models for business application experience
  • PhD, or Master's degree and 6+ years of applied research experience
  • neural deep learning methods and machine learning

Other signals

  • develop novel optimization algorithms and ML models
  • planetary scale
  • balance network utilization, resilience, cost efficiency, and latency
  • network capacity planning
  • bring algorithms from prototype to production at global scale