Senior Security Researcher and Principal Security Researcher (multiple Positions)

Microsoft Microsoft · Big Tech · Redmond, WA +1 · Security Research

This role focuses on building the evaluation backbone for safe, reliable, and efficient agentic engineering within Microsoft Security. The primary responsibility is to create systems that assess AI agents, models, prompts, tools, and orchestration patterns for production readiness. This involves designing and implementing evaluation harnesses, benchmarks, and grading pipelines, measuring various performance metrics, and integrating these evaluations into CI/CD and release decision processes. The goal is to ensure agentic systems are measurable, reproducible, governable, and continuously improving, acting as a trust system for the expansion of AI autonomy in security workflows.

What you'd actually do

  1. Design and build end-to-end evaluation harnesses for agentic security and engineering workflows, including triage, remediation, repo readiness, escalation, tool use, and scan-to-verified-closure paths.
  2. Create representative benchmark suites and golden datasets that include normal, edge, adversarial, failure-recovery, regression, and production-derived cases.
  3. Implement deterministic validators, automated graders, trace analyzers, result stores, comparison views, and workflow adapters that make evaluations repeatable and actionable.
  4. Measure task success, correctness, safety and policy compliance, failure recovery, latency, tool-call behavior, token usage, total cost per successful outcome, and human-review effort.
  5. Compare models, prompts, tools, memory strategies, policies, and orchestration patterns under consistent conditions and help teams understand quality-versus-efficiency tradeoffs.

Skills

Required

  • Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field OR Master's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 3+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection OR Bachelor's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 4+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection OR equivalent experience.

Nice to have

  • Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 3+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
  • Master's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 4+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection

What the JD emphasized

  • agentic engineering systems
  • Evals are the trust system
  • agentic security and engineering workflows
  • agentic systems measurable, reproducible, governable, and continuously improving
  • agent builders, product teams, security engineers, data scientists, program managers, and leadership

Other signals

  • building evaluation systems for agentic AI
  • creating benchmarks and datasets
  • measuring quality, safety, reliability, latency, and cost
  • integrating evaluations into CI/CD and release gates
  • connecting offline evaluation results with production feedback