Data Scientist, Amazon Connect

Amazon Amazon · Big Tech · NY +1 · Data Science

This role focuses on analyzing large datasets to understand customer behavior, identify outliers and anomalies, and improve model performance within the Amazon Connect contact center platform. The data scientist will translate business problems into data science solutions, develop mechanisms for data cleaning and outlier detection, and assess data gaps to improve feature engineering and model accuracy. The role is part of a new team building AI systems with a focus on early-stage development and significant impact.

What you'd actually do

  1. Categorizing Customer Idiosyncrasies
  2. Detecting and Cleaning Up Outliers
  3. Deep Diving Customer Issues
  4. Assessing Data Gaps

Skills

Required

  • 2+ years of data scientist experience
  • 3+ years of data querying languages (e.g. SQL), scripting languages (e.g. Python) or statistical/mathematical software (e.g. R, SAS, Matlab, etc.) experience
  • 3+ years of machine learning/statistical modeling data analysis tools and techniques, and parameters that affect their performance experience
  • Experience applying theoretical models in an applied environment

Nice to have

  • Experience in Python, Perl, or another scripting language
  • Experience in a ML or data scientist role with a large technology company
  • machine learning explainability is a plus

What the JD emphasized

  • trace decisions in data from raw data through complex models to their impact business metrics

Other signals

  • analyze data from massive data sets
  • categorize customer idiosyncrasies
  • identify outliers
  • systematically detect anomalies
  • trace decisions in data from raw data through complex models to their impact business metrics
  • translate well-defined business problems into data science problems
  • develop mechanisms to clean up outliers for downstream consumption
  • deep dive discrepancies and determine if there is an issue with our model or if we are giving the customer better results than they are used to
  • assess data gaps
  • derive features in creative ways from existing data sources
  • estimate the benefit of getting a new data stream in terms of accuracy improvement