Member of Technical Staff - Software Engineer (ai Infra)- Mai Superintelligence Team

Microsoft Microsoft · Big Tech · Zürich, ZH, Switzerland +2 · Software Engineering

Software Engineer role focused on building and optimizing AI infra for foundational models, specifically pretraining and inference on large GPU clusters. The role involves benchmarking, profiling, debugging, and tuning training and inference for generative AI, contributing to the roadmap, and ensuring reliability and performance of AI jobs. It sits within the Microsoft AI organization, working on Copilot and other consumer AI products.

What you'd actually do

  1. Develop and tune the pretraining scalable software for Nvidia GB200 72NVL CX8 and AMD MIxxx architectures.
  2. Benchmark GB200 and AMD MIxxx GPU clusters.
  3. Gather data and insights to develop the pretraining compute roadmap.
  4. Care deeply about conversational AI and its deployment.
  5. Actively contribute to the development of AI models that are powering our innovative products.

Skills

Required

  • Bachelor's Degree in Computer Science, or related technical discipline AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • Experience with generative AI.
  • Experience with distributed computing.

Nice to have

  • Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • Experience in leading technical projects and supporting architectural decisions with data.

What the JD emphasized

  • pretraining scalable software
  • training and inference infrastructures
  • generative AI
  • foundational models
  • large compute-capacity
  • supercomputer consisting of thousands of GPUs
  • reliability, runtime performance, and health
  • instrument best known state-of-the-art and novel tools and techniques
  • smooth operation of the AI jobs
  • architectural changes
  • influence roadmap of relevant software and hardware components

Other signals

  • building foundational models
  • training and inference infrastructures
  • generative AI
  • supercomputer consisting of thousands of GPUs
  • benchmark, profile, debug and tune the training and inference of generative AI