Sr Lead AI Platform Engineer

JPMorgan Chase JPMorgan Chase · Banking · GLASGOW, LANARKSHIRE, United Kingdom · Asset & Wealth Management

Sr Lead AI Platform Engineer for JPMorgan Chase's International Private Bank (IPB) Technology AIML Team. This role focuses on designing, building, and owning the platform, tooling, and infrastructure for agentic AI and ML products, moving from individual use cases to multiple production platforms. Responsibilities include deployment pipelines, model serving, containerization, orchestration, environment management, reliability, observability, security, and capacity management. The role also involves mentoring engineers and setting engineering standards.

What you'd actually do

  1. Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management
  2. Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services
  3. Builds the tooling and paved paths that let AI engineers ship agentic AI and ML products safely and quickly
  4. Implements platform-level security, secrets management, and access controls to firm-wide standards
  5. Partners with data and AI engineers to productionise models, and coordinates with external infrastructure functions to reduce operational dependency and risk

Skills

Required

  • Python
  • infrastructure-as-code
  • Kubernetes
  • containerisation
  • cloud-native deployment patterns
  • CI/CD
  • deployment and release automation
  • serving ML models
  • scaling ML models
  • monitoring ML models
  • data-intensive services in production
  • observability tooling
  • metrics
  • logging
  • tracing
  • production incident response
  • operating ML / LLM workloads in production
  • inference reliability
  • cost/performance management
  • communication skills
  • set standards other engineers adopt
  • Certified Kubernetes Application Developer (CKAD)
  • Master's degree in Computer Science, Engineering, or a related technical field (or equivalent applied experience)

Nice to have

  • Site Reliability Engineering (SRE)
  • reliability practices
  • SLOs
  • error budgets
  • financial services technology
  • JPM-internal platform, cloud, and AI/ML infrastructure

What the JD emphasized

  • production AI/ML services
  • agentic AI and ML products
  • MLOps
  • LLMOps
  • production incident response
  • operating ML / LLM workloads in production

Other signals

  • production AI/ML services
  • agentic AI and ML products
  • MLOps
  • LLMOps