Software Engineer, Systems ML (technical Leadership)

Meta Meta · Big Tech · Bellevue, WA

Seeking a principal-level Software Engineer to define technical strategy and execution for Systems ML Engineering. This role involves architecting large-scale ML infrastructure (training, inference, compilers, HPC, on-device optimization), solving complex cross-system challenges, shaping roadmaps, and leveraging AI-native workflows. The ideal candidate has deep systems expertise and can drive resolution of difficult problems at organizational scale.

What you'd actually do

  1. Identify and solve the most complex cross-system ML infrastructure challenges spanning training, inference, compiler optimization, and hardware-software co-design, including problems that have resisted prior solution attempts
  2. Define extensible architectural standards and technical foundations for ML systems that enable consistency and reliability across multiple engineering organizations
  3. Develop and own the multi-year technical roadmap for ML systems infrastructure, balancing short-term delivery with long-term platform health and competitive positioning
  4. Leverage AI-native tooling and workflows as a force multiplier to eliminate entire categories of engineering toil and accelerate cross-disciplinary work across the ML systems stack
  5. Drive performance improvements across large-scale ML training and inference systems by identifying bottlenecks that span multiple subsystems, ownership boundaries, and abstraction layers

Skills

Required

  • Software engineering
  • ML systems domains
  • AI infrastructure
  • ML compilers
  • high-performance computing
  • GPU architecture
  • ML frameworks
  • on-device optimization
  • architecting and delivering large-scale ML training or inference infrastructure
  • leading multi-year cross-functional technical initiatives
  • developing high-performance ML systems infrastructure in C++, Python, or CUDA
  • hardware-software co-design
  • influencing technical direction and engineering practices
  • integrating AI tools to optimize/redesign workflows
  • defining and operationalizing reliability, performance, and correctness standards for distributed ML training or large-scale inference systems

Nice to have

  • equivalent practical experience
  • prompt/context engineering
  • agent orchestration
  • contributing to industry-wide ML systems efforts through publications, open-source projects, or standards bodies
  • responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • MLIR
  • XLA
  • TVM

What the JD emphasized

  • 12+ years of experience in software engineering with deep specialization in one or more ML systems domains including AI infrastructure, ML compilers, high-performance computing, GPU architecture, ML frameworks, or on-device optimization
  • Experience architecting and delivering large-scale ML training or inference infrastructure that has had measurable impact across multiple engineering organizations
  • Experience leading multi-year cross-functional technical initiatives, including defining metrics, managing dependencies, and driving execution across organizational boundaries
  • Experience developing high-performance ML systems infrastructure in C++, Python, or CUDA, including work at the intersection of hardware and software
  • Experience influencing technical direction and engineering practices across multiple teams through written proposals, design reviews, and stakeholder alignment
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Experience contributing to industry-wide ML systems efforts through publications, open-source projects, or standards bodies
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Track record of applying AI tools and automation to redesign engineering workflows, with demonstrated efficiency or quality improvements at organizational scale
  • Experience with ML compiler stacks such as MLIR, XLA, or TVM, or with hardware-software co-design for custom ML accelerators
  • Experience defining and operationalizing reliability, performance, and correctness standards for distributed ML training or large-scale inference systems

Other signals

  • architectural foundations
  • large-scale machine learning infrastructure
  • training systems
  • inference pipelines
  • ML compilers
  • high-performance computing frameworks
  • on-device optimization
  • cross-system ML infrastructure challenges
  • multi-year technical roadmaps
  • AI-native workflows
  • deep systems expertise
  • organizational scale