Senior Lead Software Engineer - ML Engineer for Agent Platform

JPMorgan Chase JPMorgan Chase · Banking · Jersey City, NJ +1 · Commercial & Investment Bank

Senior Lead Software Engineer - ML Engineer for Agent Platform at JPMorgan Chase. This role involves owning the end-to-end architecture of an agent runtime platform (NEO), including secure execution, communication, memory, and evaluation. The engineer will drive technical decisions, set standards, and lead the adoption of AI-assisted engineering practices. The role requires experience building or operating LLM/agent systems in production, with a strong understanding of responsible AI use within a regulated financial services environment.

What you'd actually do

  1. Owns end-to-end architecture of the NEO agent runtime, including secure execution and isolation (micro-VMs such as Firecracker/Kata), agent-to-agent (A2A) communication, Model Context Protocol (MCP) tooling, the memory layer, and evaluation infrastructure
  2. Executes creative software solutions, design, development, and technical troubleshooting with the ability to think beyond routine or conventional approaches to build solutions or break down technical problems
  3. Develops secure and high-quality production code, and reviews and debugs code written by others; sets code and design standards adopted across teams building on NEO
  4. Drives team adoption of enterprise-authorized AI-assisted engineering practices to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team
  5. Applies and shapes the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation

Skills

Required

  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • Hands-on practical experience delivering system design, application development, testing, and operational stability for platforms, runtimes, or distributed systems in production
  • Advanced in one or more programming language(s); strong Python plus a systems language (Go or Rust) for performance-sensitive runtime work
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
  • Demonstrated experience building or operating LLM/agent systems in production, including tracing, evaluations, and guardrails
  • Proficient in all aspects of the Software Development Life Cycle
  • Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security
  • Demonstrated proficiency in software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning)
  • In-depth knowledge of the financial services industry and their IT systems
  • Practical cloud native experience; production Kubernetes expected

Nice to have

  • Experience architecting secure code execution and sandboxing with micro-VMs (Firecracker, Kata, gVisor) for multi-tenant isolation
  • Experience designing or implementing agent protocols (A2A, MCP) and multi-agent orchestration
  • Experience with agent or distributed memory systems — memory nodes, episodic/semantic memory, and graph-backed retrieval (Graph RAG)
  • Exposure to LLMs, RAG architectures, vector databases, and embedding-based retrieval systems
  • Fine-grained authorization (OpenFGA / Zanzibar-style) and policy engines (OPA/Rego)
  • Evaluation infrastructure for agents — offline/online evals, regression suites, and LLM-as-judge quality/safety gating in CI
  • Proficiency with Infrastructure as Code (Terraform) and containerized deployments (Docker, Kubernetes)

What the JD emphasized

  • secure execution and isolation
  • evaluation infrastructure
  • responsible AI use
  • building or operating LLM/agent systems in production
  • regulated environment

Other signals

  • Agent runtime platform
  • Architecture for agent execution, communication, memory, and evaluation
  • Secure execution and isolation
  • Responsible AI use in engineering workflows
  • Building or operating LLM/agent systems in production