Lead Software Engineer- Python / Pyspark / Java / Bigdata / Data Modernization / AI

JPMorgan Chase JPMorgan Chase · Banking · Jersey City, NJ +1 · Asset & Wealth Management

Lead Software Engineer role focused on building AI-first analytics and data product solutions within a global prime brokerage team. The role involves defining and enforcing an AI-driven data product lifecycle, building autonomous agents for data engineering tasks, and modernizing the strategic data mesh. Requires expertise in big data platforms, Python/PySpark/Java, and applied AI in ML pipelines, LLMs, and agentic frameworks, with a strong emphasis on governance and responsible AI use.

What you'd actually do

  1. Collaborate with business and technology teams to develop AI‑first analytics and data product solutions
  2. Define and enforce architecture for an AI‑driven data product lifecycle: semantic extraction/alignment, automated lineage and DQ, pipeline code generation, mesh registration, and self‑service consumption
  3. Build and operate autonomous agents for data engineering that detect schema drift, propose transformations, reconcile semantics, triage data incidents, and maintain governance evidence under human‑in‑the‑loop controls
  4. Design analytics platforms capable of running reporting and other analytics; explore innovative ideas by building real‑time and batch analytics solutions
  5. Establish appropriate monitoring and alerting of solution events related to performance, scalability, availability, and reliability

Skills

Required

  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • Proven leadership delivering AI‑first analytics and data engineering at scale, including productized data mesh patterns, semantic layers, and analytical data product lifecycle ownership
  • Experience developing data ingestion and integration processes, sourcing data from multiple platforms, and applying data cleansing/transformation rules for analytics-ready datasets
  • Deep hands‑on experience with big data and modern data platforms (e.g., Spark, Databricks, Snowflake, Iceberg) and building robust pipelines and data lake/lakehouse frameworks
  • Strong programming capability in Python and PySpark, or Java, with strong CI/CD and containerization practices
  • Applied AI expertise in ML pipelines, NLP/LLMs, and agentic frameworks to build autonomous agents for engineering tasks (schema drift detection, semantic reconciliation, incident triage, governance evidence generation) under human‑in‑the‑loop controls
  • Governance proficiency across lineage, data quality, and access control with evidence generation aligned to regulatory expectations (e.g., BCBS 239‑class lineage/DQ)
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices

Nice to have

  • Comfortable working in an agile and collaborative environment; strong written and verbal communication skills

What the JD emphasized

  • AI-first engineering mindset
  • AI-driven data product lifecycle
  • autonomous agents
  • human-in-the-loop
  • data mesh
  • ML/LLM solutions
  • governance
  • regulatory expectations
  • AI-assisted engineering practices
  • responsible AI use

Other signals

  • AI-driven data product lifecycle
  • autonomous agents for data engineering
  • data mesh modernization
  • ML/LLM solutions for entity resolution