Engineering Manager, Machine Learning Platform

Affirm Affirm · Fintech · United States · Remote · Checkout

Engineering Manager for Affirm's ML Training & Serving team, focusing on building and operating production ML infrastructure for model training, deployment, GPU compute, and low-latency serving. The role involves people leadership, technical guidance, and roadmap execution for ML infrastructure capabilities.

What you'd actually do

  1. Partner with senior ICs and engineering leadership to define and execute the roadmap for ML training & serving, spanning model training, deployment workflows, GPU infrastructure, and low-latency model serving.
  2. Lead, coach, and grow a team of platform engineers while staying closely engaged in technical decisions and execution.
  3. Drive delivery and operational health for the team’s infrastructure, balancing reliability, developer experience, performance, and cost.
  4. Evaluate and adopt modern ML infrastructure capabilities as Affirm’s needs evolve, including transformer-based workloads and GPU compute.
  5. Collaborate with ML modeling, product, and infrastructure teams to ensure the platform supports Affirm’s highest-priority ML initiatives.

Skills

Required

  • People leadership
  • Technical judgment in ML infrastructure
  • ML training infrastructure
  • Model serving infrastructure
  • Deployment workflows
  • GPU infrastructure
  • Production ML systems
  • Distributed systems
  • Transformer architectures
  • Large-scale training
  • Large-scale serving
  • Systems thinking
  • Platform delivery
  • Recruiting
  • Coaching engineers
  • Cross-functional collaboration

Nice to have

  • Applied ML modeling

What the JD emphasized

  • 7+ years of industry experience in software and/or machine learning engineering, including 2+ years managing engineers
  • Strong experience building and operating production ML or distributed systems infrastructure
  • hands-on experience in at least one of model training, model serving, deployment workflows, or GPU infrastructure
  • Track record of delivering platforms or infrastructure that improve the productivity and impact of engineering or ML teams.

Other signals

  • ML Platform
  • ML Training
  • Model Serving
  • GPU Infrastructure
  • Transformer Workloads