What you'd actually do

Define and drive the technical strategy for AI performance tooling and validation across multiple layers of the AI software stack, including programming models, compilers, runtimes, libraries, and model-serving APIs.

Architect scalable benchmarking and regression-detection systems for OpenAI and other state-of-the-art LLMs across GPUs, Microsoft accelerators, and emerging AI hardware platforms.

Lead deep performance investigations across model architecture, kernels, runtimes, networking, scheduling, and hardware behavior, and guide teams toward durable optimizations for production-scale training and inference.

Establish performance quality bars, readiness signals, and operating mechanisms that enable faster model and hardware bring-up while reducing regressions, deployment risk, and total hardware footprint.

Influence and align senior engineers, researchers, product leaders, and hardware partners across Microsoft and OpenAI to deliver high-impact, production-ready AI performance improvements.

Other signals

Owns inference performance of OpenAI and other state of the art large language model (LLM) models

Work directly with OpenAI on the models hosted on the Azure OpenAI service

Lead the design of benchmarking and performance tooling systems used to evaluate OpenAI and other frontier LLMs

Identify systemic bottlenecks, and translate performance insights into production-ready improvements

Reduce time-to-deploy, improve hardware efficiency, and directly support Microsoft Azure's capex goals

Overview

The Artificial Intelligence (AI) Frameworks team at Microsoft develops AI software that enables running AI models everywhere, from world’s fastest AI supercomputers, to servers, desktops, mobile phones, internet of things (IoT) devices and internet browsers. We collaborate with our hardware teams and partners, both internal and external, and operate at the intersection of AI algorithmic innovation, purpose-built AI hardware, systems, and software. We are a team of highly capable and motivated people that pride themselves on a collaborative and inclusive culture. We own inference performance of OpenAI and other state of the art large language model (LLM) models and work directly with OpenAI on the models hosted on the Azure OpenAI service serving some of the largest workloads on the planet with trillions of inferences per day in major Microsoft products, including Office, Windows, Bing, SQL Server, and Dynamics.

As a Principal Software Engineer - Performance Tooling on the team, you will set technical direction across organizations, define the architecture and engineering strategy for AI performance validation and optimization, and drive execution across hardware, runtime, compiler, and model-serving partners. You will lead the design of benchmarking and performance tooling systems used to evaluate OpenAI and other frontier LLMs across GPUs and Microsoft hardware, identify systemic bottlenecks, and translate performance insights into production-ready improvements that reduce time-to-deploy, improve hardware efficiency, and directly support Microsoft Azure's capex goals.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities

Architect scalable benchmarking and regression-detection systems for OpenAI and other state-of-the-art LLMs across GPUs, Microsoft accelerators, and emerging AI hardware platforms.

Influence and align senior engineers, researchers, product leaders, and hardware partners across Microsoft and OpenAI to deliver high-impact, production-ready AI performance improvements.

Mentor and raise the technical bar for engineers across teams by modeling engineering excellence, improving design quality, and creating reusable systems that scale beyond a single project or product line.

Qualifications

Required/Minimum Qualifications:

Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to C++, Python OR equivalent experience.

**Other Requirements: **

Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. This includes passing the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

**Preferred/Additional Qualifications: **

Master's Degree in Computer Science or related technical field AND 15+ years technical engineering experience with coding in languages including, but not limited to, C++, Python, or equivalent systems programming languages

OR Bachelor's Degree in Computer Science or related technical field AND 18+ years technical engineering experience with coding in languages including, but not limited to, C++, Python, or equivalent systems programming languages

OR equivalent experience.

8+ years of practical experience building, debugging, and optimizing high-performance distributed systems, AI inference/training workloads, or accelerator-backed compute platforms.

Deep experience with DNN/LLM inference or training performance, including one or more deep learning frameworks such as PyTorch, TensorFlow, or ONNX Runtime and accelerator programming environments such as CUDA, ROCm, Triton, or equivalent.

Recognized technical depth in software engineering, distributed systems, computer architecture, GPU/accelerator architecture, and hardware/software co-design for AI workloads.

Demonstrated ability to lead end-to-end performance analysis and optimization for state-of-the-art LLMs, HPC applications, or large-scale production AI services, including expert-level use of profiling, tracing, and observability tools.

Proven ability to influence technical strategy across organizations, resolve ambiguity, and drive alignment among senior engineering, research, product, and hardware stakeholders.

Track record of independently leading multi-team or cross-organizational initiatives from strategy through execution, with measurable impact on performance, reliability, cost efficiency, or developer productivity.

#AIInfra

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $139,900 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here: https://careers.microsoft.com/us/en/us-corporate-pay

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about **requesting accommodations.**

Overview

Responsibilities

Architect scalable benchmarking and regression-detection systems for OpenAI and other state-of-the-art LLMs across GPUs, Microsoft accelerators, and emerging AI hardware platforms.

Influence and align senior engineers, researchers, product leaders, and hardware partners across Microsoft and OpenAI to deliver high-impact, production-ready AI performance improvements.

Qualifications

Required/Minimum Qualifications:

**Other Requirements: **

**Preferred/Additional Qualifications: **

OR equivalent experience.

8+ years of practical experience building, debugging, and optimizing high-performance distributed systems, AI inference/training workloads, or accelerator-backed compute platforms.

Recognized technical depth in software engineering, distributed systems, computer architecture, GPU/accelerator architecture, and hardware/software co-design for AI workloads.

Proven ability to influence technical strategy across organizations, resolve ambiguity, and drive alignment among senior engineering, research, product, and hardware stakeholders.

#AIInfra

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here: https://careers.microsoft.com/us/en/us-corporate-pay

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Principal Software Engineer, Performance Tooling

What you'd actually do

Skills

Required

Nice to have

What the JD emphasized

Other signals