Senior And/or Principal Software Engineer - Performance

Microsoft Microsoft · Big Tech · United States · Software Engineering

This role focuses on optimizing the inference performance of large language models (LLMs) like those from OpenAI, enabling them to run efficiently across various hardware, from supercomputers to mobile devices. The engineer will benchmark, debug, and optimize model performance on GPUs and custom hardware, developing software tools and frameworks to accelerate deployment and reduce computational costs. The work is critical for large-scale AI services within Microsoft products and Azure OpenAI.

What you'd actually do

  1. Identify and drive improvements to end-to-end inference performance of OpenAI and other state-of-the-art LLMs
  2. Measure, benchmark performance on Nvidia/AMD GPUs and first party Microsoft silicon
  3. Optimize and monitor performance of LLMs and build SW tooling to enable insights into performance opportunities ranging from the model level to the systems and silicon level to improve customer experience and reduce the footprint of the computing fleet
  4. Enable fast time to market of LLMs/models and their deployments at scale by building SW tools that afford velocity in porting models on new Nvidia and AMD GPUs
  5. Design, implement, and test functions or components for our AI/DNN/LLM frameworks and tools

Skills

Required

  • Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience
  • coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • software design and development skills
  • demonstrated history of solving technical problems
  • entrepreneurial approach
  • ability to take initiative and move fast

Nice to have

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience
  • OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience
  • Technical background and solid foundation in software engineering principles, computer architecture, GPU architecture, HW neural net acceleration
  • Experience in end-to-end performance analysis and optimization of state of the art LLMs
  • proficiency using GPU profiling tools
  • Experience in DNN/LLM inference
  • experience in one or more DL frameworks such as PyTorch, Tensorflow, or ONNX Runtime
  • familiarity with CUDA, ROCm, Triton
  • Cross-team collaboration skills
  • desire to collaborate in a team of researchers and developers

What the JD emphasized

  • running AI models everywhere
  • inference performance
  • state of the art LLMs
  • trillions of inferences per day
  • benchmark performance
  • optimize performance
  • build SW tooling
  • enable fast time to market
  • deployments at scale
  • porting models
  • AI/DNN/LLM frameworks and tools
  • performance and/or efficiency
  • software design and development skills
  • solving technical problems
  • tackle the hardest problems
  • building a full end-to-end AI stack

Other signals

  • running AI models everywhere
  • inference performance of OpenAI and other state of the art LLM models
  • trillions of inferences per day
  • optimize performance
  • enable fast time to market of LLMs/models and their deployments at scale