Principal Product Manager

Microsoft Microsoft · Big Tech · Redmond, WA +1 · Product Management

Principal Product Manager to embed with applied researchers and ML engineers working on post-training (RLHF, fine-tuning, preference-tuning) for Copilot's models. This role defines 'good' for model quality, tone, and safety, runs data campaigns, builds behavior evals, and translates user feedback into training priorities. It requires a technical understanding of post-training processes and a focus on translating qualitative judgment into quantitative metrics for AI models.

What you'd actually do

  1. Embed in the post-training loop: work day-to-day with research and applied ML teams during fine-tuning and RLHF cycles, reviewing model outputs and giving structured feedback on quality, tone, and behavior
  2. Define what "good" means for Copilot's models: set the priorities for capability, personality, and safety tradeoffs across surfaces (Microsoft 365 Copilot, Copilot Studio agents, Windows) — where "good" is often genuinely ambiguous and hasn't been decided before
  3. Run RLHF/preference data campaigns: partner with data and labeling teams to scope what human preference data gets collected, for which behaviors, and why
  4. Build and maintain behavior evals: translate qualitative judgment calls ("this response felt too hedgy," "this refused when it shouldn't have") into evaluation sets that can be tracked release over release
  5. Own the feedback loop: turn user research, enterprise customer feedback, and production incident learnings into concrete post-training priorities — closing the gap between "users are unhappy with X" and "the next fine-tune addresses X"

Skills

Required

  • Bachelor's Degree AND 10+ years experience in product/service/program management or software development OR equivalent experience
  • Technical grounding in post-training (RLHF, DPO/preference optimization, instruction tuning, data curation)
  • Ability to translate qualitative judgment into concrete training priorities
  • Experience with evaluation dashboards and model outputs

Nice to have

  • 6+ years experience taking a product, feature, or experience to market
  • 8+ years experience improving product metrics for a product, feature, or experience in a market
  • 8+ years experience disrupting a market for a product, feature, or experience

What the JD emphasized

  • post-training
  • RLHF
  • fine-tuning
  • preference-tuning
  • instruction tuning
  • behavior evals
  • Responsible AI
  • enterprise requirements

Other signals

  • post-training
  • RLHF
  • fine-tuning
  • preference-tuning
  • evals
  • Copilot