Applied Scientist

Microsoft Microsoft · Big Tech · Redmond, WA +2 · Applied Sciences

The Copilot Search Multimedia Index Selection team is looking for a Data and Applied Scientist II to build the next generation platform for Bing and Microsoft AI. The role focuses on developing advanced AI and machine learning solutions for image content discovery, selection, ranking, and processing to improve the quality and relevance of multimedia products. Responsibilities include developing and evaluating multimodal models, contributing to real-time indexing systems, and working with large-scale datasets to build and optimize ML solutions for production environments.

What you'd actually do

  1. Develop expertise in relevant machine learning, deep learning, and multimodal AI techniques, and apply them to solve challenging problems in image indexing, content understanding, selection, and ranking.
  2. Partner with scientists, engineers, and product managers to understand business and product requirements, translate them into data-driven solutions, and deliver measurable product impact.
  3. Design, implement, and evaluate machine learning and deep learning models, including Large Language Models (LLMs), Small Language Models (SLMs), and multimodal models for document understanding, content discovery, parsing, clustering, selection, and ranking.
  4. Conduct large-scale data analysis and experimentation to identify opportunities for improving image index quality, freshness, relevance, and diversity.
  5. Develop robust evaluation methodologies and analyze experiment results to drive model improvements and support data-informed decision making.

Skills

Required

  • Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 2+ years related experience
  • Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 1+ year(s) related experience
  • Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field
  • equivalent experience

Nice to have

  • Solid machine learning / computer vision / natural language processing experience and results in academy or industry.
  • Publications in major ML/IR/NLP/CV conferences.
  • Experience in open source (spark, hbase, hadoop etc) or microsoft internal tech (co

What the JD emphasized

  • multimodal models
  • image content discovery
  • selection and ranking
  • near real-time indexing
  • petabyte-scale datasets

Other signals

  • multimodal models
  • image content discovery
  • selection and ranking
  • near real-time indexing
  • petabyte-scale datasets