Principal Software Engineer

Microsoft Microsoft · Big Tech · Redmond, WA +1 · Software Engineering

Principal Software Engineer role focused on optimizing the end-to-end inference performance of OpenAI and other state-of-the-art LLMs within Microsoft's Cloud & AI organization. The role involves identifying and driving improvements to key components and pipelines to enhance system performance and efficiency, with a strong emphasis on CPU/GPU optimization and leveraging profiling tools.

What you'd actually do

  1. Identify and drive improvements to end-to-end inference performance of OpenAI and other state of the art LLMs.
  2. Speeding up/reducing complexity of key components/pipelines to improve performance and/or efficiency of our systems.
  3. Creates, implements, optimizes, debugs, refactors, and reuses code to establish and improve performance and maintainability, effectiveness, and return on investment (ROI).
  4. Acts as a Designated Responsible Individual (DRI) and guides other engineers by developing and following the playbook, working on call to monitor system/product/service for degradation, downtime, or interruptions, alerting stakeholders about status and initiates actions to restore system/product/service for simple and complex problems when appropriate.
  5. Communicate and collaborate with our partners both internal and external

Skills

Required

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience
  • coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python

Nice to have

  • Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience
  • Bachelor's Degree in Computer Science or related technical field AND 15+ years technical engineering experience
  • 4+ years’ practical experience working on high performance applications and performance debug and optimization on CPU's/GPU's.
  • Technical background and solid foundation in software engineering principles, computer architecture, GPU architecture, HW neural net acceleration.
  • Experience in end-to-end performance analysis and optimization of state of the art LLMs, HPC applications including proficiency using GPU profiling tools.

What the JD emphasized

  • end-to-end inference performance
  • state of the art LLMs
  • performance debug and optimization on CPU's/GPU's
  • end-to-end performance analysis and optimization of state of the art LLMs

Other signals

  • LLM inference performance optimization
  • state of the art LLMs
  • GPU profiling tools