Meta is seeking a principal-level Software Engineer to drive technical strategy and execution across our Systems ML Engineering organization. In this role, you will define the architectural foundations that power large-scale machine learning infrastructure, spanning training systems, inference pipelines, ML compilers, high-performance computing frameworks, and on-device optimization. You will identify and solve the hardest cross-system ML infrastructure challenges, shape multi-year technical roadmaps, and amplify the impact of engineering teams through AI-native workflows and deep systems expertise. This is a role for engineers who identify problems others miss and drive them to resolution at organizational scale.
Responsibilities
Identify and solve the most complex cross-system ML infrastructure challenges spanning training, inference, compiler optimization, and hardware-software co-design, including problems that have resisted prior solution attempts Define extensible architectural standards and technical foundations for ML systems that enable consistency and reliability across multiple engineering organizations Develop and own the multi-year technical roadmap for ML systems infrastructure, balancing short-term delivery with long-term platform health and competitive positioning Leverage AI-native tooling and workflows as a force multiplier to eliminate entire categories of engineering toil and accelerate cross-disciplinary work across the ML systems stack Drive performance improvements across large-scale ML training and inference systems by identifying bottlenecks that span multiple subsystems, ownership boundaries, and abstraction layers Establish invariants, correctness proofs, and systemic reliability practices that prevent whole classes of failures across ML infrastructure pipelines Partner with research, hardware, and product engineering teams to translate theoretical ML systems advances into production infrastructure that delivers measurable efficiency and capability gains Assess emerging AI and computing technologies, evaluate competitive ML infrastructure trends, and influence organizational strategy to ensure technical competitiveness Mentor engineers across the organization by providing customized coaching, leading engineering programs, and establishing a culture of thoroughness and high craft in ML systems development Communicate complex ML systems architecture and strategy clearly to technical and non-technical audiences, producing reference-quality design documents and roadmap artifacts
Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 12+ years of experience in software engineering with deep specialization in one or more ML systems domains including AI infrastructure, ML compilers, high-performance computing, GPU architecture, ML frameworks, or on-device optimization Experience architecting and delivering large-scale ML training or inference infrastructure that has had measurable impact across multiple engineering organizations Experience leading multi-year cross-functional technical initiatives, including defining metrics, managing dependencies, and driving execution across organizational boundaries Experience developing high-performance ML systems infrastructure in C++, Python, or CUDA, including work at the intersection of hardware and software Experience influencing technical direction and engineering practices across multiple teams through written proposals, design reviews, and stakeholder alignment Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies Experience contributing to industry-wide ML systems efforts through publications, open-source projects, or standards bodies Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) Track record of applying AI tools and automation to redesign engineering workflows, with demonstrated efficiency or quality improvements at organizational scale Experience with ML compiler stacks such as MLIR, XLA, or TVM, or with hardware-software co-design for custom ML accelerators Experience defining and operationalizing reliability, performance, and correctness standards for distributed ML training or large-scale inference systems