Senior Software Engineer

Microsoft Microsoft · Big Tech · Redmond, WA +1 · Software Engineering

Senior Software Engineer role focused on defining and building agentic reliability platform capabilities for Microsoft Monitoring solutions within Azure Monitor. The role involves architecting systems for monitoring intelligence, telemetry, diagnostics, safe remediation, and operational automation, with a strong emphasis on agentic experiences and responsible AI.

What you'd actually do

  1. Define architecture for agentic reliability systems spanning monitoring, telemetry, incident management, service topology, deployment signals, and operational knowledge.
  2. Lead platform capabilities for automated detection, triage, root-cause assistance, mitigation recommendations, safe execution, and post-incident learning.
  3. Establish engineering standards for safe agentic operations, including identity, access, compliance, rollback, auditability, change management, and human escalation.
  4. Influence service teams to adopt consistent monitoring, SLOs, alert quality, incident automation, live-site readiness, and operational excellence practices.
  5. Identify high-impact reliability gaps and convert them into platform investments, architectural improvements, and reusable automation.

Skills

Required

  • Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • Ability to meet Microsoft, customer and/or government security screening requirements

Nice to have

  • Experience building production-scale platforms, cloud services, distributed systems, or reliability automation.
  • Experience architecting complex systems across service boundaries and driving execution across partner teams without direct authority.
  • Experience with observability architecture, monitoring systems, incident response, service health modeling, operational automation, and production debugging.
  • Solid judgment around production safety, automation risk, customer impact, security, privacy, compliance, and responsible AI-assisted and agentic automation.
  • Experience leading agentic automation, AI-assisted diagnostics, autonomous remediation, intelligent operations, or reliability platform efforts.
  • Experience with Azure Monitor, Log Analytics, Application Insights, Kusto/KQL, Azure Resource Graph, Azure DevOps, GitHub, or similar monitoring and observability ecosystems.
  • Experience creating organization-level reliability metrics such as SLO compliance, alert quality, time to detect, time to mitigate, human effort saved, automation coverage, and incident recurrence.
  • Experience mentoring senior engineers and defining technical strategy across monitoring, observability, incident response, and production engineering disciplines.

What the JD emphasized

  • agentic reliability platform capabilities
  • agentic operations
  • responsible AI-assisted and agentic automation

Other signals

  • agentic experience
  • agentic reliability platform capabilities
  • monitoring intelligence
  • telemetry
  • diagnostics
  • safe remediation
  • operational automation
  • automated detection
  • triage
  • root-cause assistance
  • mitigation recommendations
  • safe execution
  • post-incident learning
  • responsible AI-assisted and agentic automation