Engineering Manager, Machine Learning Platform

Affirm Affirm · Fintech · United States · Remote · Checkout

Engineering Manager for ML Platform at Affirm, leading a team focused on ML training and serving infrastructure, including model training, deployment, GPU infrastructure, and low-latency serving. The role involves people leadership, technical guidance, roadmap definition, and ensuring platform reliability to support ML priorities.

What you'd actually do

  1. Partner with senior ICs and engineering leadership to define and execute the roadmap for ML training & serving, spanning model training, deployment workflows, GPU infrastructure, and low-latency model serving.
  2. Lead, coach, and grow a team of platform engineers while staying closely engaged in technical decisions and execution.
  3. Drive delivery and operational health for the team’s infrastructure, balancing reliability, developer experience, performance, and cost.
  4. Evaluate and adopt modern ML infrastructure capabilities as Affirm’s needs evolve, including transformer-based workloads and GPU compute.
  5. Collaborate with ML modeling, product, and infrastructure teams to ensure the platform supports Affirm’s highest-priority ML initiatives.

Skills

Required

  • People management
  • Technical leadership in ML infrastructure
  • ML training infrastructure
  • Model serving infrastructure
  • GPU infrastructure
  • Production ML systems
  • Distributed systems
  • Transformer architectures
  • Deep learning
  • Recruiting and developing engineers

Nice to have

  • Applied ML modeling

What the JD emphasized

  • 7+ years of industry experience in software and/or machine learning engineering, including 2+ years managing engineers
  • Strong experience building and operating production ML or distributed systems infrastructure
  • hands-on experience in at least one of model training, model serving, deployment workflows, or GPU infrastructure
  • Solid understanding of ML data needs, including training datasets, data quality, reproducibility, and evaluation data.
  • Familiarity with modern ML workloads, such as deep learning, transformer architectures, and large-scale training or serving.

Other signals

  • ML Platform
  • ML Training
  • Model Serving
  • GPU Infrastructure
  • Transformer Workloads