Senior Engineering Manager, Model Infrastructure

Harvey Harvey · AI Frontier · San Francisco, CA · Engineering

Senior Engineering Manager for Model Infrastructure at Harvey, leading a team responsible for the platform powering all model requests. The role involves managing a team, defining the technical roadmap for reliability, scalability, and cost-efficiency, building systems for model provisioning and failover, owning the multi-provider model platform, driving the evolution of the Unified Model Controller, improving observability, supporting new model launches, improving inference efficiency, and building infrastructure for future model training efforts. The role requires significant software engineering and management experience, with a strong technical background in distributed systems and operational excellence. Experience with AI infrastructure, LLM serving, and multiple model providers is a plus.

What you'd actually do

  1. Lead and grow a high-performing team of software engineers responsible for Harvey's Model Infrastructure platform.
  2. Define the technical roadmap for model reliability, scalability, and operational excellence.
  3. Build highly reliable systems for model provisioning, capacity management, failover, and incident response across multiple AI providers.
  4. Own Harvey's multi-provider model platform, including provider integrations, SDK upgrades, API migrations, and onboarding new model providers.
  5. Drive the evolution of our Unified Model Controller (UMC) and Model Selector platform to automatically detect degraded models and intelligently route traffic based on health, latency, quality, compliance, and cost.

Skills

Required

  • 8+ years of software engineering experience
  • multiple years managing high-performing engineering teams
  • leading teams responsible for large-scale distributed systems or cloud infrastructure
  • Strong technical background that enables you to guide architectural decisions and mentor senior engineers
  • Experience operating highly available production services with strong reliability and operational excellence
  • Experience building platforms that require scalability, observability, automation, and cost optimization
  • Strong cross-functional leadership skills with the ability to partner effectively across Engineering, Research, Product, and external vendors
  • Excellent communication skills and the ability to influence technical strategy across organizations
  • A passion for building teams and developing engineering talent

Nice to have

  • Experience with AI infrastructure, LLM serving, or machine learning platforms
  • Experience working with multiple model providers such as OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source model ecosystems
  • Experience building inference platforms, model gateways, traffic routing systems, or policy-based serving infrastructure
  • Experience with Kubernetes, cloud infrastructure, distributed systems, and large-scale observability platforms
  • Experience supporting GPU infrastructure, model training platforms, or ML infrastructure
  • Familiarity with data platforms and technologies such as Spark, Kafka, Flink, Airflow, or Iceberg
  • Experience leading organizations through periods of rapid growth and technical transformation

What the JD emphasized

  • frontier agentic AI
  • enterprise-grade platform
  • scaling fast
  • operating third-party models
  • building the infrastructure that enables Harvey to train, evaluate, deploy, and operate our own frontier AI models
  • power the company's next phase of growth
  • model reliability, scalability, and operational excellence
  • multi-provider model platform
  • Unified Model Controller (UMC)
  • Model Selector platform
  • health, latency, quality, compliance, and cost
  • observability
  • token usage analytics, cost reporting, and end-to-end model telemetry
  • inference efficiency, reduce infrastructure costs, and increase model utilization
  • future model training efforts
  • data pipelines, model operations, training environments, and AI platform capabilities
  • long-term AI infrastructure strategy
  • vendor relationships
  • AI infrastructure, LLM serving, or machine learning platforms
  • multiple model providers
  • inference platforms, model gateways, traffic routing systems, or policy-based serving infrastructure
  • GPU infrastructure, model training platforms, or ML infrastructure

Other signals

  • leading teams responsible for large-scale distributed systems or cloud infrastructure
  • operating highly available production services with strong reliability and operational excellence
  • building platforms that require scalability, observability, automation, and cost optimization
  • future model training efforts, including data pipelines, model operations, training environments, and AI platform capabilities