Senior Machine Learning Engineer – Digital Intelligence

JPMorgan Chase JPMorgan Chase · Banking · Palo Alto, CA +1 · Consumer & Community Banking

Senior Machine Learning Engineer at JPMorgan Chase focused on Digital Intelligence, specializing in large language modeling, optimization, and interpretability. The role involves zero-to-one ML development, foundation model training, and implementing research into production-quality code, with a focus on financial data and recommendation systems. The position requires strong software engineering skills and expertise in LLMs and distributed training frameworks.

What you'd actually do

  1. Research and prototype next-generation architectures for structured and unstructured data
  2. Develop novel pre-training objectives tailored to financial event sequences and heterogeneous profile data
  3. Implement research ideas in production-quality code
  4. Mentor engineers on ML best practices; translate research advances into deployable systems
  5. Optimize training throughput for large data sources

Skills

Required

  • PhD with 2+ years OR Master’s degree with 4+ in Computer Science, with training and work experience in Machine Learning, LLM/NLP or similar fields
  • Deep LLM and Transformer expertise — strong command of attention mechanisms, positional encodings such as RoPE, and the ability to handle multi-modal data inputs effectively
  • PyTorch proficiency at scale — hands-on experience with distributed training frameworks including FSDP and DeepSpeed, alongside practical memory optimization techniques
  • Foundation model training — proven experience in pre-training from scratch and designing tokens and vocabularies for complex, heterogeneous data sources including tabular, temporal, and graphical formats
  • Strong software engineering skills — ability to build robust, production-quality systems that perform reliably at scale
  • Prior experience with financial data and recommendation systems

Nice to have

  • Publication record at top AI/ML venues
  • Experience optimizing serving infrastructure is a plus
  • Experience with post-training LLMs and network optimization algorithms, as well as interpretability or steering techniques for LLMs
  • Experience working with large-scale compute infrastructure
  • Experience shipping a real-world product, project, or feature
  • Experimental rigor and ablation design when benchmarking LLM optimizations
  • Strong communication and accountability skills, with a collaborative mindset and strong work ethic

What the JD emphasized

  • reliable, secure, and scalable
  • zero-to-one machine learning development
  • foundation model training
  • pre-training from scratch

Other signals

  • LLM-based methods
  • zero-to-one machine learning development
  • foundation model training
  • pre-training from scratch