Staff Site Reliability Engineer

Writer Writer · AI Frontier · New York, NY · Engineering, product & design

This role is for a Staff Site Reliability Engineer at WRITER, a company focused on enterprise generative AI. The engineer will be responsible for ensuring the availability, performance, and reliability of WRITER's AI platform. This involves using and building AI-native approaches for operational tasks, designing scalable infrastructure on public cloud providers, owning the observability stack, leading incident response, and collaborating with product and engineering teams. The role requires extensive experience in SRE/DevOps, cloud platforms, containerization, IaC, and programming languages like Python, Go, or Java.

What you'd actually do

  1. Use and build AI native approaches for operational tasks and infrastructure management and platforms using Python, Go, or similar languages, significantly reducing manual toil across our production environment
  2. Design and implement scalable, fault-tolerant infrastructure AI solutions on public cloud providers (AWS, GCP, Azure) to support WRITER's rapidly expanding, high-traffic AI platform
  3. Own the reliability, performance, and efficiency of WRITER’s core services, defining and upholding stringent Service Level Objectives (SLOs) and Error Budgets
  4. Own the observability stack for monitoring, logging, and alerting systems to ensure rapid detection of issues across our complex distributed systems
  5. Lead incident response, post-mortems, and root cause analyses, applying learnings to proactively prevent future outages and build a more resilient system architecture

Skills

Required

  • 7+ years of experience in Site reliability engineering, DevOps, Production engineering, Cloud platform or a similar role focused on building and operating large-scale, high-availability production systems
  • Deep expertise with cloud platforms (AWS strongly preferred), containerization technologies like Docker and Kubernetes, and Infrastructure-as-Code tools such as Terraform
  • Strong proficiency in programming languages such as Python, Java, Go for automation and monitoring
  • Knowledge of monitoring and logging tools (e.g., Prometheus, Grafana, ELK Stack) to maintain system health and performance
  • Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems
  • Excellent communication, collaboration, and problem-solving skills, with a talent for building strong relationships and Connecting with cross-functional teams
  • A strong sense of ownership and accountability, eager to Own mission-critical systems and drive them toward peak performance and unparalleled reliability

What the JD emphasized

  • AI native approaches
  • high-traffic AI platform
  • Own the reliability
  • Own the observability stack
  • Own mission-critical systems