AI Networking

Microsoft Microsoft · Big Tech · United States · Software Engineering

This role focuses on designing and scaling high-performance networks for AI training and inference systems. It involves architecting, implementing, and optimizing network fabrics (Ethernet, InfiniBand, ROCE) that connect thousands of GPUs, ensuring ultra-low latency and congestion-free transport for large-scale AI workloads.

What you'd actually do

  1. Advanced ROCE transport design, congestion control, ECN/WRED/DCTCP tuning
  2. Fabric architecture, topology planning, network modeling, and scaling strategy
  3. Telemetry, observability, reliability engineering, and automated troubleshooting
  4. Develop and tune the deployment of novel routing techniques to achieve reliability in large networks
  5. AI training + inference cluster bring-up, performance benchmarking, and root-cause analysis

Skills

Required

  • Bachelor's Degree in Computer Science or related technical field
  • 6+ years technical engineering experience
  • coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python

Nice to have

  • Master's Degree in Computer Science or related technical field
  • 8+ years technical engineering experience
  • 12+ years technical engineering experience

What the JD emphasized

  • scale the distributed Ethernet and InfiniBand fabrics that connect hundreds of thousands of GPUs
  • engineer ultra-low-latency ROCE networks
  • design congestion-free transport mechanisms
  • optimize lossless fabrics at 10k–100k+ GPU scale

Other signals

  • building the fabric that connects frontier-class datacenters
  • enables multi-gigawatt AI supercomputers
  • supports the training of the most sophisticated AI models
  • design, bring up, and scale the distributed Ethernet and InfiniBand fabrics that connect hundreds of thousands of GPUs
  • engineer ultra-low-latency ROCE networks
  • design congestion-free transport mechanisms
  • optimize lossless fabrics at 10k–100k+ GPU scale