Applied AI ML - Analyst

JPMorgan Chase JPMorgan Chase · Banking · Jersey City, NJ +1 · Asset & Wealth Management

This role focuses on building and maintaining end-to-end data pipelines on AWS for AI/ML models, with a specific emphasis on designing autonomous AI agents. The role involves data ingestion, transformation, curation, and serving, utilizing technologies like Spark, SQL, and Snowflake. It also requires collaboration with data science and ML teams to support model training and deployment, and includes implementing security and governance controls within a regulated financial environment.

What you'd actually do

  1. Build and maintain end-to-end batch and streaming data pipelines (ingestion → transformation → curation → serving) on AWS.
  2. Develop scalable transformations using Spark (PySpark/Scala) and SQL, optimizing for performance, reliability, and cost.
  3. Implement and optimize Snowflake solutions (schemas, tables, views, micro-partitioning/clustering considerations, warehouse sizing, query tuning, data sharing patterns where applicable).
  4. Create robust orchestration and scheduling (e.g., Airflow/MWAA, AWS Step Functions, etc.) with monitoring, alerting, retries, and operational runbooks.
  5. Collaborate with data science and ML teams to design, build, and maintain scalable data infrastructure that supports the training, deployment, and monitoring of AI/ML models.

Skills

Required

  • Bachelors degree in a quantitative or technical discipline or significant practical experience in industry.
  • Strong SQL skills (complex joins, window functions, query optimization, performance troubleshooting)
  • experience with AWS data services and cloud-native patterns (e.g., S3, IAM, KMS, Glue, Lambda, EMR, Step Functions, Kinesis/MSK—relevant mix).
  • Solid understanding of data modeling
  • Practical experience with Snowflake (data loading, transformations, optimization, access controls).
  • Proficiency in Python (or equivalent) for pipeline development, automation, and testing.
  • Experience with Git and CI/CD practices; familiarity with engineering standards for code reviews and automated testing.

Nice to have

  • Prior experience of developing solutions for Financial domain
  • Exposure to distributed model training, and deployment
  • Familiarity with techniques for model explainability and self-validation
  • Hands-on development experience with Spark (PySpark preferred) for large-scale processing.

What the JD emphasized

  • design autonomous AI agents
  • data infrastructure that supports the training, deployment, and monitoring of AI/ML models
  • security and governance controls suitable for banking

Other signals

  • design autonomous AI agents
  • data infrastructure that supports the training, deployment, and monitoring of AI/ML models
  • continuously evaluate and adopt emerging AI technologies