Senior Software Engineer Ii, Developer Experience / Operational Excellence

Samsara Samsara · Enterprise · London, United Kingdom · Platform

This role focuses on building and improving operational excellence within the Developer Experience organization. It involves designing and implementing automated reliability systems, incident management tooling, and observability infrastructure. A key aspect is contributing to AI-driven operational tooling for autonomous remediation and partnering with product engineering teams to enhance their operational posture. The role also involves defining and championing operational best practices.

What you'd actually do

  1. Design and build automated reliability and self-healing systems that protect production at scale, including automated rollbacks, deploy safeguards, and fault mitigation, and deliver them as platform tooling that engineering teams across the company adopt for their own services
  2. Own and improve incident management tooling and on-call health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden
  3. Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency
  4. Contribute to AI-driven operational tooling that goes beyond triage, building toward autonomous remediation where AI detects issues, takes corrective action, and self-recovers with minimal human involvement
  5. Drive incident prevention by identifying systemic patterns and ruthlessly eliminating operational toil. You have deep empathy for on-call engineers and a bias toward making their lives better

Skills

Required

  • Designing and building products in a software engineering team
  • infrastructure and/or platform engineering
  • Observability and reliability
  • operational metrics and data analysis
  • architecting monitoring frameworks
  • SLO platforms
  • automated response workflows
  • large-scale enterprise software applications
  • Developer Experience (DevEx) & Internal Tooling

Nice to have

  • Datadog (or equivalent observabilty tooling like New Relic, Grafana)

What the JD emphasized

  • AI-driven operational tooling
  • autonomous remediation
  • automated safeguards
  • observability tooling
  • incident management tooling
  • operational excellence best practices
  • automated rollbacks
  • deploy safeguards
  • fault mitigation
  • alert noise
  • actionable signals
  • system health
  • latency
  • corrective action
  • self-recovers
  • systemic patterns
  • operational toil

Other signals

  • AI-driven operational tooling
  • autonomous remediation
  • automated safeguards
  • observability tooling
  • incident management tooling