Senior Sre - Platform (managed Kubernetes Infrastructure)

Elastic Elastic · Enterprise · Canada · Platform - SRE

This role is for a Senior Site Reliability Engineer focused on building, scaling, and maturing the multi-cloud platform for hosting internal and external services, specifically managing Kubernetes infrastructure at scale. The role involves automating system engineering efforts, developing software and tooling for infrastructure, and ensuring the reliability of global infrastructure.

What you'd actually do

  1. Taking an engineering approach in leading technical initiatives for automating system engineering efforts to guarantee the reliability of the global Elastic infrastructure.
  2. Growing our global Platform infrastructure to meet the increasing scaling demands by developing and maintaining software, tooling and automations.
  3. Collaborating in an environment with an inclusive approach, and focusing on operational excellence, and uplifting others.
  4. Responding to and preventing repeated customer impact in response to major incidents and prioritised problem management.

Skills

Required

  • Production experience in Public Cloud Service Providers
  • managing Kubernetes infrastructure at scale
  • software engineering background
  • Golang

Nice to have

  • SaaS product in a public cloud
  • Infrastructure-as-Code tooling such as Crossplane or Terraform
  • Kubernetes-at-scale infrastructure across multiple cloud providers
  • containerized services (such as Docker)
  • leading and improving alerting and major incident management standard processes metrics systems (e.g. Elastic Stack, Prometheus, Influx)
  • system administration with professional skills in Linux on distributed systems at scale
  • diagnosed or designed, implemented and created solutions with the Elastic Stack
  • thriving in a self-organizing and sharing in a globally distributed team environment
  • coaching and mentoring