Principal Product Manager

Microsoft Microsoft · Big Tech · Redmond, WA +1 · Product Management

Principal Product Manager to own the evaluation systems for Copilot, determining readiness for launch based on quality, safety, and task success metrics. This role involves designing and building evaluation suites, gating model and feature launches, and translating model behavior into product requirements.

What you'd actually do

  1. Own the eval lifecycle end-to-end: design, build, and maintain evaluation suites (offline benchmarks, live A/B experiments, agentic task suites, red-team/adversarial sets) that measure quality, safety, and task success across Copilot surfaces (Microsoft 365, Windows, Edge, Copilot Studio agents, etc.)
  2. Gate model and feature launches: define entry/exit criteria for shipping model updates or new capabilities, and make the ship/no-ship call in partnership with research and engineering leadership
  3. Translate model behavior into product requirements: read outputs and failure modes directly, identify patterns, and turn them into prioritized fixes — whether that's a training data gap, a prompt/system-message change, or a UX guardrail
  4. Build feedback loops: stand up pipelines that turn user feedback, support signals, and enterprise customer escalations into structured eval cases, so regressions get caught before they reach customers
  5. Partner cross-functionally: work daily with applied scientists, ML engineers, Responsible AI, and Trust & Safety teams to represent user needs in model development, and represent model capabilities/limitations back to the broader product org

Skills

Required

  • Bachelor's Degree AND 8+ years experience in product/service/program management or software development OR equivalent experience

Nice to have

  • Bachelor's Degree AND 12+ years experience in product/service/program management or software development OR equivalent experience
  • 4+ years experience taking a product, feature, or experience to market
  • 6+ years experience improving product metrics for a product, feature, or experience in a market
  • 6+ years experience disrupting a market for a product, feature, or experience

What the JD emphasized

  • own how we measure that
  • build and run the evaluation systems
  • determine whether a model update, a new Copilot skill, or an agentic workflow is ready to ship
  • make the call on whether something is ready for customers
  • Own the eval lifecycle end-to-end
  • Gate model and feature launches
  • make the ship/no-ship call
  • define what "good" looks like for ambiguous, high-stakes scenarios

Other signals

  • evaluation systems
  • measure quality, safety, and task success
  • gate model and feature launches
  • define entry/exit criteria for shipping
  • make the ship/no-ship call