Software Engineer, Core Machine Learning

Meta Meta · Big Tech · Sunnyvale, CA +1

Staff Software Engineer on the Core Machine Learning team, focused on building and scaling foundational ML infrastructure and systems for training, model serving, and feature engineering. The role involves architecting and delivering high-impact ML platform capabilities to enable thousands of engineers and researchers to build and ship state-of-the-art ML models at scale.

What you'd actually do

  1. Architect and own large-scale ML infrastructure systems, including distributed training frameworks, model serving platforms, and feature computation pipelines that support production workloads across Meta's product surface
  2. Lead the technical design and implementation of foundational ML platform components, evaluating trade-offs across performance, reliability, and developer experience
  3. Drive end-to-end delivery of major ML infrastructure initiatives, coordinating across teams and disciplines to align on priorities, manage dependencies, and execute phased rollouts
  4. Identify and resolve performance bottlenecks in ML training and inference systems through instrumentation, profiling, and targeted optimization
  5. Define and enforce service level objectives for core ML platform services, building dashboards, alerting systems, and runbooks to reduce mean time to mitigation during incidents

Skills

Required

  • Software engineering
  • Machine learning systems
  • ML infrastructure
  • Large-scale distributed systems
  • Distributed training frameworks
  • Model serving infrastructure
  • Feature engineering pipelines
  • Performance analysis and optimization
  • Instrumentation
  • Profiling
  • Bottleneck resolution
  • Service level objectives (SLOs)
  • Alerting systems
  • Runbooks
  • Incident management
  • Testing strategies
  • Safe rollout patterns
  • AI-accelerated development workflows
  • Technical design
  • Architectural decisions
  • Data analysis
  • Design reviews
  • Debugging complex distributed system issues
  • Technical roadmap contribution
  • Prompt/context engineering
  • Agent orchestration
  • Emerging AI technologies
  • AI tools to optimize/redesign workflows
  • AI-assisted development tools
  • Code generation
  • Automated testing
  • Intelligent debugging
  • Hardware-software co-design for ML workloads
  • Quantization
  • Model compression
  • Resource-efficient AI techniques
  • Responsible, ethical AI practices
  • Risk assessment
  • Bias mitigation
  • Quality and accuracy reviews
  • PyTorch
  • TensorFlow
  • Ray

Nice to have

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • Building or contributing to open-source ML frameworks or platform tooling

What the JD emphasized

  • 8+ years of experience in software engineering with a focus on machine learning systems, ML infrastructure, or large-scale distributed systems
  • Experience designing and implementing production ML systems such as distributed training frameworks, model serving infrastructure, or large-scale feature engineering pipelines
  • Experience leading major technical initiatives from design through production, including cross-team coordination and phased rollout management
  • Experience with performance analysis and optimization of ML training or inference workloads, including profiling, instrumentation, and bottleneck resolution

Other signals

  • building and scaling foundational ML infrastructure
  • architect and deliver high-impact ML platform capabilities
  • enable thousands of engineers and researchers across Meta to build and ship state-of-the-art machine learning models at scale