Principal Security Research Manager

Microsoft Microsoft · Big Tech · Redmond, WA +1 · Security Research

This role focuses on building the evaluation backbone for safe, reliable, and efficient agentic engineering within Microsoft Security. The primary responsibility is to create systems that assess AI agents, models, prompts, tools, and orchestration patterns for readiness in production security workflows. This involves designing evaluation harnesses, creating benchmark suites, implementing validators and graders, measuring various performance aspects (success, safety, cost, latency), and integrating these evaluations into CI/CD, release gates, and decision processes. The goal is to make agentic systems measurable, reproducible, governable, and continuously improving, ensuring autonomy can safely expand and releases meet quality and safety bars.

What you'd actually do

  1. Design and build end-to-end evaluation harnesses for agentic security and engineering workflows, including triage, remediation, repo readiness, escalation, tool use, and scan-to-verified-closure paths.
  2. Create representative benchmark suites and golden datasets that include normal, edge, adversarial, failure-recovery, regression, and production-derived cases.
  3. Implement deterministic validators, automated graders, trace analyzers, result stores, comparison views, and workflow adapters that make evaluations repeatable and actionable.
  4. Measure task success, correctness, safety and policy compliance, failure recovery, latency, tool-call behavior, token usage, total cost per successful outcome, and human-review effort.
  5. Integrate evaluations into engineering workflows, CI/CD, release gates, and decision processes so material agent changes are supported by reproducible evidence before production rollout.

Skills

Required

  • Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 3+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection OR Master's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 4+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection OR Bachelor's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 6+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection OR equivalent experience.
  • 1+ year(s) people management experience.

Nice to have

  • Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 5+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
  • Proven software engineering experience building production systems, developer platforms, test infrastructure, automation frameworks, data pipelines, quality systems, or reliability tooling.
  • Ability to design and implement evaluation systems end to end, including task definition, dataset creation, harness implementation, scoring, analysis, and operational integration.

What the JD emphasized

  • evaluation backbone for safe, reliable, and efficient agentic engineering
  • systems that determine when an AI agent, model, prompt, tool, memory strategy, or orchestration pattern is ready to be used in production security and engineering workflows
  • Evals are the trust system for that shift.
  • make agentic systems measurable, reproducible, governable, and continuously improving
  • build the common evaluation platform and methodology used by MSec agent programs
  • design and implement evaluation systems end to end

Other signals

  • building evaluation systems for agentic engineering
  • creating benchmarks and golden datasets
  • integrating evaluations into CI/CD and release gates
  • connecting offline evaluation results with production feedback