Technical Program Manager, Compute Qualification

Together AI Together AI · Data AI · San Francisco, CA · Business Operations

This role is responsible for qualifying new compute capacity for Together AI, ensuring it meets technical standards for training and inference workloads. The Technical Program Manager will manage the end-to-end qualification process, coordinating with engineering partners, reviewing specifications, and making go/no-go recommendations. The role requires strong technical understanding of data center infrastructure and program management skills.

What you'd actually do

  1. Own and continuously improve the end-to-end qualification process for new compute capacity, from initial provider intake through final go/no-go recommendation.
  2. Run multiple provider evaluations in parallel, setting timelines, tracking status, and keeping every stakeholder aligned on what is needed and by when.
  3. Partner with infrastructure engineering, network engineering, data center engineering, and SRE teams to plan and coordinate technical validation, then translate their findings into clear decisions for leadership.
  4. Review provider technical specifications and questionnaire responses for completeness and accuracy, flagging gaps, inconsistencies, and risks that warrant follow-up.
  5. Conduct first-pass analysis of provider data yourself: compare specifications across suppliers , sanity-check performance claims, and surface issues before deeper engineering review.

Skills

Required

  • technical program or project management
  • infrastructure program management
  • large-scale compute environments
  • organization and stakeholder management
  • server and GPU hardware
  • high-performance networking (InfiniBand or Ethernet fabrics)
  • storage
  • power and cooling fundamentals
  • Python
  • SQL
  • written and verbal communication

Nice to have

  • qualifying, commissioning, or accepting GPU clusters or HPC infrastructure
  • AI training and inference infrastructure
  • interconnect topologies
  • cluster bring-up
  • acceptance testing
  • AI/HPC cluster design
  • hardware vendors
  • colocation providers
  • cloud capacity providers

What the JD emphasized

  • critical hardware performance metrics
  • training or inference workloads