Principal Software Engineer

Microsoft Microsoft · Big Tech · Redmond, WA +1 · Software Engineering

Principal Software Engineer to scale the backbone infrastructure for AI-powered vulnerability discovery and proofing in Windows, handling AI evals, throughput, reliability, compute, capacity, orchestration, and cost efficiency at Windows scale.

What you'd actually do

  1. AI Evals & Benchmarks: Building AI evals and benchmarks to ensure that the system stays healthy and does not regress across multiple scenarios.
  2. Throughput & Reliability: Scaling the pipeline by an order of magnitude and improve reliability.
  3. Compute, Capacity & Orchestration: Managing compute and AI infra at Windows-wide scale; smoothing capacity spikes across model-hosting backends; managing queuing/scheduling and pool prioritization infrastructure
  4. Cost & Efficiency: Driving billing/consumption estimation and cost-of-goods modeling.

Skills

Required

  • 6+ years building large-scale distributed systems / service backends (queuing, scheduling, autoscaling, multi-tenant compute)
  • Deep experience with capacity planning, throughput/latency SLAs, and cost-of-goods on cloud infrastructure (Azure preferred)
  • Track record owning architecture of a complex, ambiguous system end to end, and mentoring senior engineers
  • Comfortable negotiating and holding a hard dependency line with partner teams
  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python, OR equivalent experience.

Nice to have

  • Experience building or operating AI agent / LLM-based systems at production scale.
  • Hands-on experience with Azure DevOps (ADO) pipelines and CI/CD automation.
  • Deep experience with Microsoft Azure cloud infrastructure (compute, capacity planning, cost-of-goods).
  • Comfortable working with C# and Microsoft technologies, including hands-on experience with Microsoft Copilot CLI, Azure DevOps (ADO) pipelines and CI/CD automation.
  • Demonstrated experience designing and operating large-scale distributed systems or cloud service backends (orchestration, scheduling, autoscaling, or multi-tenant compute).

What the JD emphasized

  • scale vulnerability discovery
  • AI-scale volume
  • order of magnitude

Other signals

  • AI-powered discovery
  • autonomous systems
  • AI-scale volume
  • order of magnitude scaling