Machine Learning Engineer

Workday Workday · Enterprise · Vancouver, BC

Machine Learning Engineer on the AI Platform Context and Retrievals team, developing user experiences using Agentic AI, LLMs, and RAG. The role involves delivering ML solutions across Workday's product ecosystem, enabling training, deployment, and lifecycle management of ML models, and developing/deploying APIs/services using Docker/Kubernetes at scale. Focus on information retrieval and graph-based recommendation systems.

What you'd actually do

  1. Own exploration, design and implementation of features for our sophisticated ML platforms, pipelines and services.
  2. Be responsible for evaluation, scalability and observability of these features.
  3. Apply machine learning techniques including LLMs and natural language understanding to analyze large sets of HR and Finance-related text data, and design and launch pioneering cloud-based machine learning architectures
  4. Stay up to date with advancements in AI, LLMs, RAG, autonomous agents and orchestration frameworks to drive innovation
  5. Serve as a technical role model for more junior engineers

Skills

Required

  • Python
  • software engineering
  • cloud computing platforms (e.g. AWS, GCP, etc.)
  • Docker
  • Kubernetes
  • PySpark
  • Pytorch
  • TensorFlow
  • Sklearn
  • Pandas

Nice to have

  • Master’s or PhD degree
  • building information retrieval systems
  • graph-based recommendation systems
  • developing large language models (LLMs)
  • text generation models
  • graph-based machine learning models
  • model fine-tuning
  • model deployment
  • model evaluation
  • building services to host machine learning models in production at scale
  • data engineering
  • data wrangling
  • statistical analysis
  • unsupervised and supervised machine learning algorithms
  • natural language processing
  • information retrieval
  • recommendation system use cases
  • technically leading team

What the JD emphasized

  • taking products through applied research, design, implementation, production, and production-based evaluation
  • shipping production code and models
  • building information retrieval systems
  • developing large language models (LLMs), text generation models, or graph-based machine learning models for production
  • building services to host machine learning models in production at scale
  • technically leading team

Other signals

  • AI Platform
  • Agentic AI
  • LLMs
  • RAG
  • ML models
  • enterprise-scale systems