Senior And/or Principal Software Engineer - AI Frameworks

Microsoft Microsoft · Big Tech · United States · Software Engineering

Develops software, performance systems, and engineering tools for state-of-the-art AI models at cloud scale, focusing on model onboarding, performance, reliability, and deployment efficiency across various hardware platforms. This role involves hands-on individual contribution to solve complex systems problems related to AI workloads, measurement, and cross-organizational collaboration.

What you'd actually do

  1. Design, implement, test, and operate production-quality components across AI frameworks, runtimes, benchmarking systems, performance tooling, and service integrations.
  2. Benchmark, profile, debug, and optimize large language model training and inference workloads across GPUs and Microsoft hardware.
  3. Build automation and observability that detect regressions, improve reproducibility, surface actionable insights, and accelerate model and hardware onboarding.
  4. Drive scoped projects from problem definition through deployment, balancing delivery speed, maintainability, reliability, and measurable customer or capacity impact.
  5. Partner with researchers, model teams, infrastructure owners, and hardware vendors to diagnose cross-stack issues and deliver production-ready solutions.

Skills

Required

  • Bachelor’s Degree in Computer Science or a related technical field and 4+ years of technical engineering experience coding in languages such as C++, or Python, or equivalent experience.

Nice to have

  • Experience building and operating complex software systems
  • practical knowledge of performance analysis, benchmarking, automation, or developer tooling
  • familiarity with AI/ML frameworks such as PyTorch, TensorFlow, or ONNX Runtime
  • familiarity with GPU software and profiling technologies such as CUDA, ROCm, Triton, or equivalent
  • demonstrated cross-team collaboration and technical ownership
  • Extensive experience designing and shipping complex, high-performance or distributed software systems
  • strong foundation in software architecture, computer architecture, and accelerator-aware optimization
  • experience with AI/ML workloads and frameworks
  • demonstrated leadership of cross-team technical initiatives from strategy

What the JD emphasized

  • production-quality components
  • large language model training and inference workloads
  • automation and observability
  • model and hardware onboarding
  • cross-stack issues
  • production-ready solutions
  • technical vision
  • architecture
  • multi-release strategy
  • critical AI framework
  • performance
  • benchmarking
  • developer-productivity capabilities
  • cross-stack investigations
  • models
  • frameworks
  • compilers
  • runtimes
  • systems
  • services
  • silicon
  • common measurement
  • automation
  • observability
  • engineering mechanisms
  • scalable platform capabilities
  • model onboarding velocity
  • runtime performance
  • reliability
  • hardware utilization
  • Azure capacity efficiency
  • architecture and priorities
  • clear decisions
  • execution plans
  • hands-on technical leadership
  • prototypes
  • critical-path implementation
  • design and code reviews
  • complex debugging
  • operational readiness
  • engineering bar
  • mentoring senior engineers
  • developing technical leaders
  • advancing standards
  • quality
  • maintainability
  • inclusive collaboration
  • extensive experience designing and shipping complex, high-performance or distributed software systems
  • strong foundation in software architecture, computer architecture, and accelerator-aware optimization
  • experience with AI/ML workloads and frameworks
  • demonstrated leadership of cross-team technical initiatives from strategy

Other signals

  • AI frameworks
  • performance systems
  • engineering tools
  • cloud scale
  • model onboarding
  • performance and reliability
  • deployment time
  • hardware footprint
  • performance insights
  • platform capabilities
  • end-to-end systems problems
  • AI workloads
  • disciplined measurement
  • production impact
  • large language model training and inference workloads
  • automation and observability
  • reproducibility
  • actionable insights
  • model and hardware onboarding
  • customer or capacity impact
  • cross-stack issues
  • production-ready solutions
  • technical design reviews
  • engineering standards
  • operational health
  • mentoring
  • technical vision
  • architecture
  • multi-release strategy
  • critical AI framework
  • performance
  • benchmarking
  • developer-productivity capabilities
  • cross-stack investigations
  • models
  • frameworks
  • compilers
  • runtimes
  • systems
  • services
  • silicon
  • common measurement
  • automation
  • observability
  • engineering mechanisms
  • scalable platform capabilities
  • model onboarding velocity
  • runtime performance
  • reliability
  • hardware utilization
  • Azure capacity efficiency
  • architecture and priorities
  • clear decisions
  • execution plans
  • hands-on technical leadership
  • prototypes
  • critical-path implementation
  • design and code reviews
  • complex debugging
  • operational readiness
  • engineering bar
  • mentoring senior engineers
  • developing technical leaders
  • advancing standards
  • quality
  • maintainability
  • inclusive collaboration
  • C++
  • Python
  • complex software systems
  • performance analysis
  • benchmarking
  • automation
  • developer tooling
  • AI/ML frameworks
  • PyTorch
  • TensorFlow
  • ONNX Runtime
  • GPU software
  • profiling technologies
  • CUDA
  • ROCm
  • Triton
  • cross-team collaboration
  • technical ownership
  • shipping complex, high-performance or distributed software systems
  • software architecture
  • computer architecture
  • accelerator-aware optimization
  • AI/ML workloads
  • leadership of cross-team technical initiatives
  • strategy