Principal Associate, Data Scientist

Capital One Capital One · Banking · McLean, VA +2

This role focuses on building and deploying machine learning models for fraud prevention in the fintech domain. It involves the full lifecycle of model development, from design and training to evaluation, validation, implementation, monitoring, and continuous deployment in a production environment. The role emphasizes traditional ML techniques but also has an appetite for AI-based development, with a focus on optimizing models for challenging segments and ensuring regulatory compliance.

What you'd actually do

  1. Partner with a cross-functional team of data scientists, software engineers, and product managers to deliver a product customers love
  2. Leverage a broad stack of technologies — Python, Conda, AWS, H2O, Spark, and more — to reveal the insights hidden within huge volumes of numeric and textual data
  3. Build machine learning models through all phases of development, from design through training, evaluation, validation, and implementation, monitoring, and supporting continuous model deployment and maintenance in a production environment.
  4. Collaborate on the design and maintenance of production data science solutions, including writing clear technical documentation and ensuring models adhere to software development best practices.
  5. Manage model risk and maintain regulatory compliance across the model lifecycle, which includes maintaining model inventory records, executing model testing and change control protocols, and collaborating on independent model validation and compliance risk assessments.

Skills

Required

  • Bachelor's Degree in a quantitative field plus 5 years of experience performing data analytics
  • Master's Degree in a quantitative field or an MBA with a quantitative concentration plus 3 years of experience performing data analytics
  • PhD in a quantitative field
  • Python
  • Spark
  • AWS
  • SQL
  • machine learning
  • model risk governance

Nice to have

  • Master’s Degree in “STEM” field plus 3 years of experience in data analytics
  • PhD in “STEM” field
  • 1 year of experience working with AWS
  • 3 years’ experience in Python, Scala, or R
  • 3 years’ experience with machine learning
  • 3 years’ experience with SQL
  • big data and distributed computing, using Spark or another comparable framework
  • technically leading and developing a team
  • traditional machine learning and emerging GenAI techniques

What the JD emphasized

  • mission-critical machine learning models
  • optimize models for highly challenging and expanding segments
  • build and deploy machine learning models
  • traditional ML model development, not GenAI-only experience
  • model risk governance

Other signals

  • build and deploy mission-critical machine learning models
  • optimizing models for highly challenging and expanding segments
  • detects and mitigates first-party fraud by building and deploying machine learning models
  • combining deep experience in traditional ML with an appetite for AI-based development
  • Build machine learning models through all phases of development, from design through training, evaluation, validation, and implementation, monitoring, and supporting continuous model deployment and maintenance in a production environment.