Advanced Data Scientist

Honeywell Honeywell · Industrial · Raleigh, NC +1

Seeking an Advanced Data Scientist to develop data-driven solutions for the connected utilities platform, focusing on predictive grid reliability and operational intelligence. The role involves designing and implementing statistical and ML models, building ML pipelines on Databricks, and applying both classical statistical methods and modern ML techniques. Responsibilities include predictive maintenance, translating business problems into modeling tasks, implementing model monitoring, and contributing to forecasting and outage prediction systems. Requires strong programming skills in Python/R/SQL, experience with time-series modeling, and familiarity with Databricks ML ecosystem and distributed computing.

What you'd actually do

  1. Design and implement statistical and machine learning models for time-series forecasting, anomaly detection, and asset health scoring across utility networks.
  2. Build and maintain end-to-end ML pipelines on Databricks from feature engineering and model training to validation, deployment, and monitoring in production.
  3. Apply classical statistical methods (GLMs, GAMs, mixed-effects models, Bayesian inference) alongside modern ML techniques (ensemble approaches, network analysis, neural networks) to solve grid operations problems.
  4. Develop predictive maintenance and degradation models for utility infrastructure using telemetry and SCADA data at scale.
  5. Translate ambiguous business problems into well-defined modeling problems with appropriate statistical frameworks - e.g., knowing when a LM/GLM is sufficient and when gradient boosting or deep learning is warranted.

Skills

Required

  • Python (PySpark, pandas, NumPy, scikit-learn, statsmodels, XGBoost)
  • R (tidyverse, lme4, glmmTMB, glmnet, mgcv)
  • SQL
  • time-series modeling (ARIMA, state-space models, LSTM, Darts, or similar)
  • Databricks ML ecosystem (Feature Store, Experiment Track, Model Serving, Mosaic AI)
  • MLflow
  • distributed computing concepts - PySpark, Optuna/Ray, Spark SQL, partitioning strategies, and medallion architecture
  • software engineering principles - version control (Git), testing, CI/CD for ML systems
  • communication of complex statistical/ML concepts to non-technical stakeholders

Nice to have

  • Master’s or PhD in Statistics, Applied Mathematics, Physics, Engineering, or a related quantitative field
  • electric or gas utility industry experience
  • AMI, SCADA, GIS, or OMS data familiarity
  • deep learning (PyTorch)
  • geospatial analysis (GeoPandas, network/graph-based modeling)
  • generative AI and LLM-based solutions
  • invention disclosure and patent process

What the JD emphasized

  • track record of deployed production models
  • production engineering skills
  • deployed production models

Other signals

  • end-to-end ML pipelines
  • production models
  • model monitoring, drift detection, and automated retraining