Senior AI Product Manager, Leaderboard

Scale AI Scale AI · Data AI · New York, NY +1 · Gen AI Product

Senior AI Product Manager to own and scale Scale's SEAL Leaderboard portfolio, defining strategy, roadmap, and operational excellence for evaluation products, benchmarks, and leaderboards. The role involves transforming model evaluations into industry benchmarks, working with AI Product Management, ML Researchers, Engineering, Operations, and Go-To-Market teams. Responsibilities include driving benchmark innovation, governance, infrastructure, customer adoption, and business impact, acting as a thought leader in AI evaluation and measurement.

What you'd actually do

  1. Own the roadmap and strategy for Scale’s SEAL Leaderboard portfolio, defining priorities across benchmark development, leaderboard launches, infrastructure investments, and product expansion.
  2. Facilitate the Leaderboard Steering Committee and drive alignment across AI-PM, ML, Engineering, Operations, and GTM stakeholders.
  3. Evaluate, prioritize, and operationalize new leaderboard proposals, ensuring alignment with customer demand, market opportunities, and company strategy.
  4. Define and manage the end-to-end leaderboard product lifecycle, from ideation and benchmark design to launch, growth, maintenance, and sunset decisions.
  5. Partner with ML researchers and domain experts to develop trustworthy evaluation methodologies, benchmark specifications, and leaderboard scoring frameworks.

Skills

Required

  • 5+ years of experience in product management, technical program management, consulting, or customer-facing technical roles.
  • Strong technical fluency, including familiarity with machine learning systems, AI model evaluation, benchmarking, or data products.
  • Experience building and scaling products that require coordination across engineering, operations, and business teams.
  • Excellent stakeholder management and executive communication skills, with a demonstrated ability to drive alignment across cross-functional organizations.
  • Strong analytical skills and the ability to translate ambiguous market and customer signals into clear product strategy.

Nice to have

  • Experience working with AI researchers, ML teams, or evaluation frameworks is strongly preferred.
  • Entrepreneurial mindset with a track record of creating new products, programs, or business initiatives from the ground up.
  • Bias for action and comfort operating in fast-moving, ambiguous environments.
  • Passion for advancing trustworthy AI evaluation and helping define industry standards for measuring frontier model capabilities.

What the JD emphasized

  • AI model evaluation
  • benchmarking
  • evaluation methodologies
  • benchmark design
  • evaluation needs
  • trustworthy AI evaluation
  • measuring frontier model capabilities

Other signals

  • Own the roadmap and strategy for Scale’s SEAL Leaderboard portfolio
  • Define the strategy, roadmap, and operational excellence of Scale’s evaluation products, benchmarks, and public/private leaderboards
  • Transform cutting-edge model evaluations into trusted industry benchmarks that influence model development and purchasing decisions