Sr Lead Infrastructure Engineer- Devops/aws

JPMorgan Chase JPMorgan Chase · Banking · GLASGOW, LANARKSHIRE, United Kingdom · Asset & Wealth Management

Lead Infrastructure Engineer (DevOps/SRE) for an AI/ML team within a financial services company. The role focuses on owning and evolving CI/CD pipelines, reliability practices, observability, and infrastructure automation for agentic AI and machine learning products. This includes managing production incident response, security, and mentoring junior engineers, with a strong emphasis on operating ML/LLM workloads in a regulated environment.

What you'd actually do

  1. Owns and evolves the team's CI/CD pipelines, release automation, and deployment tooling
  2. Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review
  3. Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting)
  4. Automates infrastructure provisioning and configuration through infrastructure-as-code
  5. Partners with platform, data, and AI engineers to harden services for production and reduce dependency on external functions

Skills

Required

  • Advanced proficiency with infrastructure-as-code (e.g., Terraform) and scripting in Python and/or shell
  • Deep hands-on experience with CI/CD tooling and building release automation at scale
  • Strong experience with Kubernetes, containerisation, and cloud-native operations
  • Proven experience running production services: observability, on-call, incident response, and reliability engineering
  • Understanding of production security and change-management controls
  • Strong communication skills and the ability to set operational standards across a team
  • Formal SRE experience in a regulated or high-availability environment
  • Master's degree in Computer Science, Engineering, or a related technical field (or equivalent applied experience)

Nice to have

  • Experience operating ML / LLM workloads in production (MLOps, inference reliability, cost/performance management)
  • Experience within financial services technology
  • Familiarity with JPM-internal platform, cloud, and observability tooling for internal candidates

What the JD emphasized

  • agentic AI and machine learning products
  • AI/ML services
  • ML / LLM workloads
  • regulated or high-availability environment
  • financial services technology

Other signals

  • Owns and evolves the team's CI/CD pipelines, release automation, and deployment tooling
  • Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review
  • Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting)
  • Automates infrastructure provisioning and configuration through infrastructure-as-code
  • Mentors junior engineers on DevOps and reliability practices and sets standards through review
  • Formal SRE experience in a regulated or high-availability environment
  • Experience operating ML / LLM workloads in production (MLOps, inference reliability, cost/performance management)