Applied AI Scientist

Microsoft Microsoft · Big Tech · Zürich, ZH, Switzerland · Applied Sciences

The Spatial AI Lab is looking for an Applied AI Scientist to train and improve multimodal AI models for new types of agents. This role involves research in ML algorithms, pre/post-training of foundational multimodal models, building scalable data and learning solutions, curating datasets, optimizing models for hardware, and collaborating with research teams. There are opportunities for publication in top-tier venues and integration into Microsoft products.

What you'd actually do

  1. Research novel machine learning algorithms and models.
  2. Work on pre and/or post training of foundational multimodal models.
  3. Build data and learning solutions for scalability, efficiency, and performance.
  4. Curate training and evaluation datasets/benchmark.
  5. Optimize models for CPUs, GPUs and NPUs and integrate into products.

Skills

Required

  • PhD in Machine Learning / Computer Vision or 3+ years of relevant industry experience
  • Engineering skills in programming languages such as Python and/or C++
  • Hands-on experience with modern deep learning frameworks (e.g. Pytorch/Tensorflow/Jax)
  • Self-motivated team-player, problem solver, and keen to learn
  • Ability to present complex technical concepts to a diverse audience

Nice to have

  • Multimodal Models hands-on experience
  • Pre and/or post training of large vision language models
  • Experience in techniques such as pruning, distillation and finetuning
  • LLMs; Large vision-language models (VLMs)
  • Video generative models and diffusion algorithms
  • Action-based transformers and Vision Language Action models (VLAs)
  • Large-Scale ML Systems Experience with large scale machine learning compute systems

What the JD emphasized

  • Publications Track record of impact, either via research publications at top-tier machine learning or computer vision conferences (NeurIPS, ICML, CVPR, ECCV, ICCV ), or via contributions to successful industry initiatives.

Other signals

  • training multimodal AI models
  • pre and/or post training of foundational multimodal models
  • optimize models for CPUs, GPUs and NPUs and integrate into products