Principal Software Engineer

Microsoft Microsoft · Big Tech · United States · Software Engineering

Principal Software Engineer role focused on optimizing GPU inference efficiency, utilization, and cost for Azure's AI infrastructure. The role involves designing, building, and operating cloud-scale systems to improve inference workloads, translating research into production services, and driving technical strategy for inference-serving systems.

What you'd actually do

  1. Design, implement, and operate cloud-scale services that improve GPU inference efficiency, utilization, and cost across Azure.
  2. Partner with Azure Systems Research and Azure AI infrastructure teams to translate algorithms, models, and systems research into production-ready infrastructure.
  3. Build telemetry, experimentation, and control-plane capabilities that identify inference-efficiency opportunities and safely validate improvements at fleet scale.
  4. Write high-quality production code in languages such as C#, Rust, Python, or C++ with engineering practices for reliability, security, and maintainability.
  5. Drive architecture and technical strategy for workload placement, capacity management, performance isolation, and optimization of inference-serving systems.

Skills

Required

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • Ability to meet Microsoft, customer and/or government security screening requirements

Nice to have

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • 4+ years of cloud infrastructure or distributed systems development experience.
  • 2+ years of experience developing, deploying, or optimizing machine learning, artificial intelligence, or inference-serving systems at scale.
  • 4+ years of experience analyzing and troubleshooting large-scale distributed systems in critical production online service environments, including experience with GPU infrastructure, machine learning inference systems, workload scheduling, capacity optimization, or performance engineering.

What the JD emphasized

  • GPU inference efficiency
  • inference cost
  • utilization
  • serving capacity
  • cloud-scale systems
  • distributed systems
  • machine learning infrastructure
  • inference-serving systems

Other signals

  • GPU efficiency
  • inference cost
  • utilization
  • serving capacity
  • cloud-scale systems
  • distributed systems
  • machine learning infrastructure