Site Reliability Engineer 2

Microsoft Microsoft · Big Tech · Hyderabad, TS, IN · Site Reliability Engineering

Site Reliability Engineer focused on building AI-driven agents and automation for incident triage and response within Microsoft's Azure Data engineering team, specifically for Power BI, Fabric, and Power Query services. The role involves on-call duties, developing agentic systems for alert ingestion, correlation, and assessment, and automating incident lifecycle processes.

What you'd actually do

  1. Incident triage and first-line response: Provide on-call coverage for incoming incidents across CDI services. Perform initial investigation, severity assessment, and routing to owning engineering teams.
  2. Agentic triage system development: Build and extend AI-driven agents that ingest ICM alerts, correlate with recent deployments and feature flag rollouts, check known-issue databases, and produce initial assessments with suggested severity and owning team.
  3. TSG and known-issue matching: Develop automation that matches incoming incidents to relevant Troubleshooting Guides (TSGs) and known issues across Fabric and Power Platform — reducing investigation time and enabling faster resolution.
  4. Auto-routing and classification: Configure and extend ICM routing rules and build intelligent classification systems based on service tree, alert signatures, and historical patterns.
  5. Incident lifecycle automation: Build agents for incident summarization, customer communications drafting, postmortem generation, and reporting, replacing manual authoring with AI-assisted workflows requiring human judgment only for high-severity incidents.

Skills

Required

  • 4+ years of software engineering experience in site reliability, Live site operations, or incident management for cloud services.
  • Good programming skills in one or more of: C#, PowerShell, Python, KQL/Kusto.
  • Experience with incident management systems and workflows (ICM, PagerDuty, ServiceNow, or similar).
  • Experience with monitoring, alerting, and observability systems (Kusto, Geneva, Grafana, or similar).
  • Ability to work in an on-call rotation across time zones in a geographically distributed team.
  • Experience interface with engineers, leadership, support, and customers.

Nice to have

  • Experience building AI/ML-driven automation, agents, or intelligent workflows (e.g., using LLMs, Copilot extensibility, MCP servers, or agentic frameworks).
  • Familiarity with Live site ecosystem management (including log traversal, incident management, telemetry analysis, etc.)
  • Experience with Azure, Power BI, and Fabric services.
  • Experience with Troubleshooting Guide (TSG) authoring and incident pattern analysis.
  • Understanding of SLA management, customer communications, and escalation workflows for cloud services.

What the JD emphasized

  • build and extend AI-driven agents
  • intelligent classification systems
  • AI-assisted workflows

Other signals

  • AI-driven agents
  • intelligent automation
  • incident triage and response
  • live site engineering