Principal Group Engineering Manager

Microsoft Microsoft · Big Tech · Bengaluru, KA, IN · Software Engineering

This role is for a Principal Group Engineering Manager at Microsoft's CoreAI team, focusing on building and evolving the AI platform for large-scale reinforcement learning and post-training of cutting-edge LLMs. The role involves technical leadership in distributed systems, ML systems, AI infrastructure, and compute orchestration to improve iteration loops, debug complex interactions, and optimize performance for mission-critical AI workloads.

What you'd actually do

  1. Build and evolve distributed services that underpin massive training runs
  2. Improve iteration loops for researchers and engineers
  3. Debug complex interactions between models and hardware
  4. Apply advanced performance optimization techniques that directly impact product quality and operational excellence
  5. Develop deep expertise in ML systems, AI infrastructure, and compute orchestration
  6. Ship platform capabilities that enable mission‑critical AI workloads for customers around the world

Skills

Required

  • Technical leadership
  • Distributed systems
  • ML systems
  • AI infrastructure
  • Compute orchestration
  • Performance optimization
  • Debugging complex interactions
  • Responsible AI practices
  • Software development lifecycle (SDLC)
  • System architecture
  • Scalability
  • Resiliency
  • Disaster recovery
  • Cost of goods sold (COGS)
  • Security
  • Privacy
  • Compliance requirements

Nice to have

  • GenAI tooling (e.g., GitHub CoPilot)

What the JD emphasized

  • build the AI platform that trains the world’s most advanced models
  • large‑scale reinforcement learning
  • large‑scale post-training
  • mission‑critical AI workloads

Other signals

  • build the AI platform that trains the world’s most advanced models
  • building next-generation systems for large‑scale reinforcement learning
  • power the full lifecycle of cutting‑edge LLMs
  • making large‑scale post-training and reinforcement learning workflows faster, safer, and more reliable
  • Build and evolve distributed services that underpin massive training runs
  • Ship platform capabilities that enable mission‑critical AI workloads