Principal Site Reliability Engineer

UiPath UiPath · Enterprise · Bangalore, India · Engineering

UiPath is seeking a Principal Site Reliability Engineer to build intelligent reliability platforms and tooling that leverage AI/ML to improve service reliability, reduce operational toil, and accelerate incident response. The role involves designing self-healing mechanisms, AI-assisted debugging tools, predictive reliability models, and AI-powered incident response systems for large-scale, cloud-native systems.

What you'd actually do

  1. Design and implement self-healing mechanisms including automated remediation workflows and intelligent retry and fallback strategies.
  2. Build internal systems that enable engineering teams to debug faster using AI-assisted tooling and proactively identify and mitigate reliability risks.
  3. Define and evolve reliability strategy using predictive reliability models(Capacity, Failure forecasting, Reliability scoring) and embed intelligent reliability practices across the engineering teams.
  4. Build AI-powered systems that determine impact and use historical data to improve detection and response over time.
  5. Influence standards for building AI-driven tooling, mentor junior and senior engineers, and elevate reliability focus across the organization.

Skills

Required

  • 7+ years of experience in SRE, Platform, Cloud infrastructure engineering roles
  • Strong conceptual understanding of distributed systems, performance bottlenecks, failure modes, and trade-offs inherent to large-scale systems.
  • Experience building applications or internal tools using LLMs to automate non-trivial workflows
  • Hands-on experience with building Agents/Copilots using modern ML frameworks (PyTorch, vLLM or equivalent) in production setting.
  • Proficiency in at least one programming language (e.g., Python, Go, or similar).
  • Experience with Infrastructure as Code (e.g., Terraform, Pulumi)
  • Experience with container orchestration (e.g., Kubernetes).
  • Hands-on experience working with one or more major cloud providers (Azure, AWS, GCP)
  • Proven experience with monitoring/observability stacks (metrics, logs, traces)
  • Experience participating in and improving incident response, blameless postmortems, and implementing systemic fixes rather than symptomatic patches.
  • Ability to partner with product, infrastructure, and engineering teams to influence architecture and reliability practices without direct authority.

Nice to have

  • practical knowledge of networking, deployments, and scaling.
  • building meaningful dashboards and alerts that improve reliability signals.

What the JD emphasized

  • AI/ML
  • AI-driven tooling
  • LLMs
  • Agents/Copilots
  • reliability

Other signals

  • building intelligent reliability platforms
  • leverage AI/ML to improve reliability
  • reduce operational toil
  • accelerate incident response
  • predictive reliability
  • self-healing capabilities
  • AI-assisted tooling
  • LLMs to automate non-trivial workflows
  • Agents/Copilots using modern ML frameworks