Sr. Applied Scientist, Silicon and Systems Group Edge Ai, Edge AI Platform

Amazon Amazon · Big Tech · CA, BC +1 · Applied Science

This role focuses on developing and validating novel evaluation methods for multimodal language models and agents, specifically for edge devices. The scientist will analyze datasets, interpret model failures, and collaborate with training teams to enhance model capabilities, with a strong emphasis on reliability and automated evaluation frameworks.

What you'd actually do

  1. Collaborate with cross-functional engineers and scientists to advance the state of the art in multimodal model evaluations for devices, including audio, images, and videos
  2. Invent and validate reliability for novel automated evaluation methods for perception tasks, such as fine-tuned LLM-as-judge
  3. Develop and extend our evaluation framework(s) to support expanding capabilities for multimodal language models
  4. Analyze large offline and online datasets to understand model gaps, develop methods to interpret model failures, and collaborate with training teams to enhance model capabilities for product use cases
  5. Work closely with other scientists, compiler engineers, data collection, and product teams to advance evaluation methods.

Skills

Required

  • building machine learning models for business application experience
  • PhD, or Master's degree and 6+ years of applied research experience
  • Experience programming in Java, C++, Python or related language
  • Experience with neural deep learning methods and machine learning

Nice to have

  • Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.
  • Knowledge of programming languages such as C/C++, Python, Java or Perl
  • Experience with Linux/Unix operating systems
  • Experience in professional software and systems development

What the JD emphasized

  • develop new evaluation methods
  • novel automated evaluation methods
  • evaluate multimodal language models and agents

Other signals

  • develop new evaluation methods for multimodal language models and agents
  • invent and validate reliability for novel automated evaluation methods
  • develop and extend our evaluation framework(s)
  • analyze large offline and online datasets to understand model gaps
  • develop methods to interpret model failures