Principal Software Engineering Manager

Microsoft Microsoft · Big Tech · United States · Software Engineering

This role manages a team building foundational components of Azure's AI networking infrastructure, focusing on high-performance, scalable, and observable networking systems for massive distributed AI training and inference. It operates at the intersection of AI, cloud infrastructure, and high-performance networking.

What you'd actually do

  1. Hire, manage, and grow a high-performing team of software engineers, fostering a culture of excellence, inclusion, and innovation.
  2. Lead the design and development of large-scale distributed systems and services that power Azure’s AI infrastructure.
  3. Drive engineering planning and execution while ensuring alignment with organizational OKRs and long-term strategy.
  4. Establish lean, scalable, and efficient processes that promote innovation and engineering rigor.
  5. Deliver best-in-class engineering by ensuring services and components are modular, secure, reliable, diagnosable, observable, and reusable.

Skills

Required

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • Ability to meet Microsoft, customer and/or government security screening requirements

Nice to have

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • 4+ years people management experience.
  • 1+ years experience building and operating networking infrastructure for hyperscale datacenters or AI clusters.
  • 1+ years hands-on experience with networking technologies in AI-specific hardware (e.g., InfiniBand, ROCE, MRC, NVLink, UALink).
  • 10+ years of professional software design and development experience in large-scale distributed systems.

What the JD emphasized

  • networking infrastructure for hyperscale datacenters or AI clusters
  • networking technologies in AI-specific hardware

Other signals

  • AI infrastructure
  • distributed training systems
  • networking infrastructure
  • large-scale model training and inference