Senior Site Reliability Engineer- Remote

ClickHouse ClickHouse · Data AI · Canada +1 · Engineering

ClickHouse is seeking a Senior Site Reliability Engineer to build and lead processes for ensuring the reliability, availability, scalability, and performance of their cloud infrastructure. This role involves collaborating with engineering teams, establishing SLOs/SLAs, ensuring monitoring and alerting, enhancing incident response, and driving chaos initiatives. The ideal candidate has a strong background in SRE, cloud platforms (AWS, Azure, GCP), distributed databases (especially ClickHouse), container orchestration (Kubernetes), and automation tools.

What you'd actually do

  1. Collaborate with various engineering teams in ClickHouse to design and implement scalable, secure, and highly available systems for ClickHouse.
  2. Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.
  3. Ensure all the infrastructure components in ClickHouse Cloud (including Dataplane, Control Plane and ClickHouse Core) have monitoring and alerting in place to ensure timely detection and resolution of incidents.
  4. Enhance and refine incident response processes and post-mortem analysis for any outages in ClickHouse Cloud including working with the support team to communicate to the impacted customers.
  5. Continuously improve the reliability and performance of our ClickHouse services.

Skills

Required

  • Site Reliability Engineering
  • Go
  • Python
  • AWS
  • Azure
  • Google Cloud Platform
  • Kubernetes
  • Docker Swarm
  • Ansible
  • Terraform
  • Puppet
  • Distributed databases
  • SQL
  • Problem solving
  • Production debugging

Nice to have

  • ClickHouse

What the JD emphasized

  • At least 8 years of experience in Site Reliability Engineering or a related field.
  • Previous experience using ClickHouse in production.
  • Hands on experience with Go and/or Python.
  • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • Excellent understanding of distributed databases and SQL, particularly ClickHouse is a major plus.
  • Hands on experience with container orchestration tools such as Kubernetes or Docker Swarm.
  • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.