Lead Software Engineer - ML Engineer for Agent Platform

JPMorgan Chase JPMorgan Chase · Banking · Jersey City, NJ +1 · Commercial & Investment Bank

Lead Software Engineer for an ML Engineer role focused on building and operating an agent runtime platform (NEO) for Payments Technology. Responsibilities include hands-on engineering of core components like secure execution, communication, memory, retrieval, and evaluation, as well as driving adoption of AI-assisted engineering practices and ensuring responsible AI use. Requires strong Python, experience with LLM/agentic systems, and knowledge of the financial services industry.

What you'd actually do

  1. Builds and operates major NEO runtime components — agent execution and sandboxing (micro-VMs), A2A and MCP integrations, the memory layer (memory nodes), retrieval, and evaluation harnesses
  2. Develops secure and high-quality production code, and reviews and debugs code written by others
  3. Drives team adoption of enterprise-authorized AI-assisted engineering practices to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team
  4. Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation
  5. Implements permission-aware, auditable execution for agents, including fine-grained authorization and runtime policy checks

Skills

Required

  • software engineering concepts
  • system design
  • application development
  • testing
  • operational stability
  • Python
  • AI-assisted software development tools
  • responsible AI use
  • LLM-power or agentic systems
  • tracing
  • evaluations
  • guardrails
  • Software Development Life Cycle
  • agile methodologies
  • CI/CD
  • Application Resiliency
  • Security
  • cloud native experience
  • production Kubernetes

Nice to have

  • LLMs
  • RAG architectures
  • vector databases
  • embedding-based retrieval systems
  • Graph RAG
  • agent protocols (A2A, MCP)
  • multi-agent orchestration
  • sandboxed/secure code execution
  • containers
  • micro-VMs
  • Firecracker
  • Kata
  • gVisor
  • agent memory
  • memory nodes
  • episodic/semantic memory
  • graph-backed retrieval
  • building or running evals for LLM/agent systems
  • Infrastructure as Code (Terraform)
  • containerized deployments
  • Docker
  • Kubernetes
  • data observability
  • quality
  • metadata management tools

What the JD emphasized

  • strong Python required
  • Hands-on experience building LLM-power or agentic systems, including tracing, evaluations, and guardrails
  • In-depth knowledge of the financial services industry and their IT systems

Other signals

  • agent runtime platform
  • agent execution
  • agent-to-agent communication
  • LLM-power or agentic systems