Applied Scientist, Amazon Music - Catalog Quality

Amazon Amazon · Big Tech · IN, KA, Bengaluru · Applied Science

The Music Catalog Quality team at Amazon Music is looking for an Applied Scientist to develop solutions for ensuring and improving the quality of catalog metadata and content. This role involves using Generative AI, classical ML, NLP, Computer Vision, and automated data validation pipelines to detect, measure, and remediate quality issues in music metadata. The scientist will own the design and development of end-to-end systems, create technical roadmaps, and drive production-level projects, working closely with scientists, engineers, and product managers.

What you'd actually do

  1. Collaborate with scientists, engineers, and product managers to define and frame business problems as ML or optimization tasks.
  2. Use machine learning, deep learning, LLMs and Agentic AI techniques to create scalable solutions for business problems
  3. Analyze and extract relevant information from large amounts of Amazon's data to help automate and optimize key processes
  4. Design, development and evaluation of AI models for predictive learning
  5. Research and implement novel machine learning and statistical approaches

Skills

Required

  • building models for business application experience
  • PhD, or Master's degree and 4+ years of CS, CE, ML or related field experience
  • Experience programming in Java, C++, Python or related language
  • Experience in any of the following areas: algorithms and data structures, parsing, numerical optimization, data mining, parallel and distributed computing, high-performance computing

Nice to have

  • Experience using Unix/Linux
  • Experience in professional software development

What the JD emphasized

  • end-to-end systems
  • production level projects
  • scalable solutions
  • predictive learning
  • novel machine learning and statistical approaches

Other signals

  • Generative AI
  • classical ML
  • Natural Language Processing
  • Computer Vision
  • automated data validation pipelines