Senior Lead Software Engineer - LLM Ops Platform Reliability

JPMorgan Chase JPMorgan Chase · Banking · GLASGOW, LANARKSHIRE, United Kingdom · Corporate Sector

This role focuses on building and operating reliable, scalable, and cost-efficient large language model (LLM) serving infrastructure. It involves designing, developing, and deploying LLM inference platforms on cloud and on-premises environments, implementing deep observability, performance tuning (quantization, parallelism), and leading reliability engineering efforts including incident response. The role also touches on building multi-agent systems and integrating AI-assisted engineering practices.

What you'd actually do

  1. Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure
  2. Build backend services and APIs that enable reliable operation of AI infrastructure in production environments
  3. Operate and scale large language model serving infrastructure, including model hosting, request routing, continuous batching, and cache optimization
  4. Deploy, host, and lifecycle-manage open-source and proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines
  5. Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads

Skills

Required

  • system design
  • application development
  • testing
  • operational stability in production environments
  • Python
  • production-grade services and tooling
  • automation
  • continuous delivery methods
  • cloud infrastructure platforms
  • infrastructure-as-code tooling
  • site reliability engineering practices
  • incident management
  • root-cause analysis
  • runbooks
  • reliability patterns
  • observability and instrumentation
  • metrics, logs, and traces
  • Kubernetes
  • container-based orchestration platforms
  • managed cloud variants
  • hosting and serving large language models on cloud-based infrastructure
  • local GPU environments
  • large language model reliability and risk considerations
  • latency and throughput trade-offs
  • model versioning
  • prompt and response logging
  • safe rollout patterns
  • enterprise-authorized AI-assisted software development tools
  • critically evaluate, validate, and refine AI-generated outputs
  • responsible AI use in engineering workflows
  • data sensitivity considerations
  • secure handling of inputs/outputs
  • resiliency and security expectations

Nice to have

  • operating large language model inference servers such as vLLM and llm-d
  • developing generative AI applications
  • AI agents
  • vector search
  • retrieval-augmented generation patterns
  • building AI agents using orchestration frameworks such as LangChain, LangGraph, CrewAI
  • operating or integrating model serving platforms such as KServe, Ray Serve, or NVIDIA Triton Inference Server
  • Amazon SageMaker JumpStart
  • SageMaker Endpoints
  • Amazon Bedrock
  • online large language model quality monitoring
  • hallucination detection
  • toxicity filtering
  • drift detection
  • open telemetry conventions
  • Contributions to open-source large language model serving or inference projects

What the JD emphasized

  • large language model serving infrastructure
  • production at scale
  • deep observability
  • cost-aware performance tuning
  • reliability, performance, and cost-efficiency
  • large language model inference platform end to end
  • operate large language model serving stacks in production at scale
  • deep instrumentation and strong operational rigor
  • large language model endpoints
  • multi-agent systems with strong orchestration
  • AI-assisted engineering practices

Other signals

  • LLM Ops Platform Reliability
  • LLM serving infrastructure
  • production at scale
  • observability
  • cost-aware performance tuning
  • reliability, performance, and cost-efficiency of the large language model inference platform end to end