Principal AI Network Hardware Systems Engineer

Microsoft Microsoft · Big Tech · Redmond, WA +4 · Hardware Engineering

The Principal AI Network Hardware Systems Engineer will lead the architecture, bring-up, validation, optimization, and deployment of networking infrastructure for Microsoft's MAIA AI platform. This role involves working across the networking stack, including high-speed SerDes, optics, cables, NICs, PHYs, switch silicon, AI communication frameworks, and distributed training systems, to deliver industry-leading AI performance and reliability for large-scale AI training and inference clusters.

What you'd actually do

  1. Define and develop networking requirements for large-scale AI training and inference clusters.
  2. Collaborate with silicon, system software, firmware, hardware, and Azure infrastructure teams to deliver scalable networking solutions from concept through datacenter deployment.
  3. Lead design and validation of IP-based AI networking solutions spanning TCP/IP, UDP, routing, congestion management, flow control, QoS, and traffic engineering.
  4. Design, validate, and optimize RDMA-based networking solutions for AI clusters.
  5. Develop and execute networking validation strategies covering functionality, performance, scale, interoperability, resiliency, and reliability.

Skills

Required

  • Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 7+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 8+ years technical engineering experience OR equivalent experience
  • 8+ years of experience in NW HW development
  • 8+ years of experience in GPU based SU/SO development
  • 8+ years of hands on experience with HS interface architecture and development

Nice to have

  • Experience with RDMA technologies, AI fabrics, and distributed training environments.
  • Understanding of RoCE, congestion control, ECN, PFC, DCQCN, and related AI networking technologies.
  • Experience with AI/ML workload communication patterns and collective operations.

What the JD emphasized

  • AI-native silicon
  • AI training and inference at hyperscale
  • MAIA AI platform
  • AI networking infrastructure
  • AI performance and reliability
  • AI communication frameworks
  • distributed training systems
  • AI training and inference clusters
  • AI networking roadmaps
  • AI infrastructure
  • AI networking solutions
  • AI fabric performance
  • AI traffic patterns
  • AI training and inference workloads
  • AI fabric
  • distributed training environments
  • AI networking technologies
  • AI/ML workload communication patterns

Other signals

  • AI-native silicon and system-level solutions
  • AI training and inference at hyperscale
  • networking infrastructure for Microsoft's MAIA AI platform
  • industry-leading AI performance and reliability