Senior Software Engineer

Microsoft Microsoft · Big Tech · Redmond, WA +1 · Software Engineering

Senior Software Engineer on the Fleet Health Team within Microsoft 365 Core Platform, responsible for improving the reliability, availability, efficiency, and operational safety of the infrastructure powering Microsoft 365 services. The role involves building and operating large-scale distributed systems that analyze telemetry from servers, storage, networking, and repair workflows, using AI/ML for proactive failure detection, predictive maintenance, automated remediation, and data-driven capacity decisions. The engineer will also build AI-assisted experiences for incident investigation and automation workflows.

What you'd actually do

  1. Design and build large-scale distributed services that improve fleet reliability, hardware health, and operational efficiency.
  2. Develop telemetry and analytics platforms that process and analyze infrastructure health signals at hyperscale.
  3. Build predictive models and intelligent services for hardware failure detection, repair recommendation, anomaly detection, and fleet risk forecasting.
  4. Analyze telemetry from servers, storage platforms, networking equipment, rack infrastructure, and datacenter systems to identify opportunities for improving reliability and availability.
  5. Build AI-assisted experiences that accelerate incident investigation, root cause analysis, and repair decision-making.

Skills

Required

  • Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python

Nice to have

  • Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • Experience designing and operating distributed systems and cloud services at scale.
  • Experience working with hardware infrastructure, storage systems, server platforms, networking systems, or datacenter operations.
  • Experience using data science, statistics, machine learning, forecasting, anomaly detection, or predictive analytics to solve engineering problems.
  • Experience with telemetry and data platforms such as Azure Data Explorer (Kusto), Spark, Fabric, Databricks, or similar analytics technologies.
  • Experience developing AI-powered operational tools, intelligent automation systems, or agent-based solutions.
  • Experience with hardware reliability engineering, fleet management, capacity planning, or infrastructure health monitoring.
  • Experience working with M365 components like Exchange, Substrate, SharePoint to improve performance, availability and supportability of services.
  • Demonstrated ability to independently drive complex technical projects from concept through production deployment.
  • Collaboration and communication skills with the ability to influence across organizations.
  • Tier 2 or Tier 3 United States Government clearance to work in secure Microsoft cloud environments.

What the JD emphasized

  • large-scale distributed systems
  • analyze telemetry
  • predictive models
  • AI and machine learning techniques
  • AI-assisted experiences
  • safe automation and remediation workflows

Other signals

  • build predictive models
  • apply AI and machine learning techniques
  • build AI-assisted experiences