Principal Data Engineer

Caterpillar Caterpillar · Industrial · Peoria, IL

Principal Data Engineer for Caterpillar's Legal, Security and Public Policy team, focusing on building scalable data pipelines and AI/ML-driven predictive analytics applications to provide insights and mitigate enterprise risks. The role involves designing and maintaining Snowflake-based data platforms, operationalizing models, and ensuring data governance.

What you'd actually do

  1. Design, build, and maintain scalable Snowflake-based data platform supporting analytics, reporting, and AI/ML workloads
  2. Identify, design, and implement internal process improvements: automating manual processes, optimizing data delivery, re-designing data pipelines for greater scalability
  3. Perform debugging, troubleshooting, modification and testing of integration solutions
  4. Operationalize developed jobs, processes and models
  5. Create solutions and methods to monitor systems and solutions

Skills

Required

  • Python
  • SQL
  • Snowflake
  • ETL/ELT design
  • Data modeling
  • Change data capture
  • Indexing
  • Partitioning
  • Query optimization
  • Cloud Platforms (Snowflake, EMR, Glue, Snap Logic)
  • Medallion architecture
  • Data lake architecture
  • Lake house architecture
  • Github
  • CI/CD
  • Communication skills
  • Interpersonal skills
  • Collaboration skills

Nice to have

  • MS in Computer Science, Computer/Data Engineering, Data Science
  • Agent assisted programming
  • Machine Learning (ML)
  • Artificial Intelligence (AI)
  • Agentic AI solutions
  • Cortex
  • Snowpipe
  • Snowpark
  • Snowsight
  • AWS EMR
  • AWS Glue
  • AWS Lambda
  • Bedrock/Agent core
  • Dataiku
  • Trade data
  • Compliance data
  • PowerBI
  • Power Apps
  • streamlit

What the JD emphasized

  • BS in Computer Science, Computer Engineering, Data Engineering, Data Science, or related field
  • 5+ years development experience with modern development languages and frameworks, preferably Python, SQL, etc.
  • Experience managing a Snowflake database/environment
  • Experience building ETL/ELT pipelines using industry platforms (Snowflake, EMR, Glue, Snap Logic or similar)
  • Experience building medallion, data lake, or lake house architectures
  • Experience with version control in Github and devops practices, like continuous integration and continuous delivery (CI/CD)
  • Good communication, interpersonal, and collaboration skills
  • Ability to communicate difficult or technical concepts in a relatable manner
  • Ability to collaborate and prioritize cross functionally to achieve results
  • Ability to guide and mentor other in coding and design best practice
  • Experience in agent assisted programming
  • Experience building and deploying machine learning (ML), artificial intelligence (AI) and/or agentic AI solutions in a business context
  • Deep Snowflake expertise including Cortex, Snowpipe, Snowpark, Snowsight, etc.
  • Deep experience in AWS EMR, Glue, Lambda, or similar
  • Experience with modern agentic AI platforms (Cortex, Bedrock/Agent core, Dataiku, etc.)
  • Experience with trade and/or compliance data

Other signals

  • building scalable data pipelines
  • data driven predictive analytics applications
  • AI/ML models
  • actionable recommendations and insights
  • enterprise risks