Infrastructure Engineer (uk)

Writer Writer · AI Frontier · London, United Kingdom · Engineering, product & design

Infrastructure engineer responsible for the availability, performance, and reliability of WRITER's AI platform, focusing on building resilient systems, automating across the stack, and championing reliability best practices. The role involves working with cloud infrastructure (AWS, GCP, Azure), Kubernetes, Terraform, and AI tooling, with a strong emphasis on using AI agents in daily workflows for incident response, automation, and code generation.

What you'd actually do

  1. Bring deep focus to one problem at a time, with the breadth to move between SRE, DevOps, Infrastructure, and Platform work over a quarter or two as the leverage shifts.
  2. Challenge the status quo and remove toil before adding features — automate operational tasks and infrastructure management with Python or Go, reject tools that don't fit the problem, and treat manual on-call work as a defect to be designed out, not a status quo to be staffed up.
  3. Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently across Kubernetes, Helm, Terraform, and the supporting cloud and AI tooling that backs WRITER's high-traffic platform.
  4. Run agents in your daily loop — Claude Code, Droid, Codex, internal skills — to investigate incidents, draft Terraform / Helm changes, write runbooks, scaffold tooling, and review PRs.
  5. Lead incident response, post-mortems, and root-cause analyses — trace failures to the underlying problem (never the symptom), apply the learning back into the architecture, and prevent the same incident from happening twice.

Skills

Required

  • 5+ years of experience in infrastructure engineering, DevOps, or a similar role
  • building and operating large-scale, high-availability production systems
  • running containerisation in production (a real cluster, not a lab)
  • experience in Helm and Terraform or Pulumi on at least one major cloud (AWS preferred)
  • good proficiency in Python or Go for automation and tooling
  • AI is part of how you ship, not a thing you've read about
  • agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop
  • built or adopted AI-assisted workflows others now use
  • strong opinions on where it's unreliable
  • Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems
  • reason from constraints and failure modes (not analogy or vendor defaults)
  • name the tradeoff in busines

Nice to have

  • GCP
  • Azure

What the JD emphasized

  • AI is part of how you ship, not a thing you've read about
  • Candidates whose actual daily workflow does not already include AI tooling will not be advanced
  • Challenge the status quo

Other signals

  • building and deploying AI agents
  • enterprise-grade LLMs
  • enterprise generative AI
  • high-traffic platform
  • AI in workflow