Senior Mlops Engineer

Oura Oura · Consumer · United States · Remote · Data Engineering & Analytics

Senior MLOps Engineer responsible for designing and evolving platforms, workflows, and governance practices for machine learning teams to develop, train, deploy, and operate ML systems reliably at scale. Focuses on MLOps platform and ML enablement, including environment, platform capabilities, operational foundations, training, orchestration, deployment, governance, experiment tracking, model packaging, and promotion processes. Collaborates with data scientists and engineers to improve the end-to-end ML lifecycle and identify/resolve issues in ML infrastructure. Also contributes to CI/CD, observability, and workflow orchestration for ML training and batch inference.

What you'd actually do

  1. Lead the development and improvement of the environment, platform capabilities, and operational foundations that support machine learning workflows across Oura.
  2. Drive the design and unification of workflows and tooling that support reliable training, orchestration, and deployment of ML systems.
  3. Partner with data scientists and engineers to improve the end-to-end ML lifecycle, from experimentation and training through deployment and governance, while making production ML systems easier to manage and maintain.
  4. Define and evolve model governance practices, including reproducibility, lineage, access controls, and operational standards.
  5. Support and help standardize ML tooling and workflows such as experiment tracking, model packaging, and promotion processes, with MLflow or similar tooling as a key part of the stack.

Skills

Required

  • 5+ years of experience in MLOps, machine learning engineering, platform engineering, data engineering, or a closely related field.
  • Hands-on experience running production workloads in AWS and a strong understanding of cloud infrastructure concepts.
  • Strong understanding of the machine learning lifecycle, including training workflows, deployment patterns across ML systems, monitoring, and ongoing model maintenance.
  • Familiarity with data science ways of working and the tooling that supports experimentation and model operations, such as MLflow.
  • Familiarity with workflow orchestration, infrastructure-as-code, and CI/CD practices for ML or data platforms.
  • Familiarity with secure access patterns, governance controls, and shared cloud or data platform services that support ML work at scale.
  • Experience building or supporting production-grade ML workflows with a focus on reliability, reproducibility, and maintainability.
  • Proven ability to drive standards and improvements across multiple teams and business domains.
  • Strong communication and collaboration skills, with the ability to work effectively with both technical and non-technical stakeholders.
  • Comfort operating in a distributed team with a high degree of ownership and autonomy.

Nice to have

  • Experience supporting multiple data science teams through shared MLOps platforms or tooling.
  • Prior experience with Databricks.
  • Experience with governance and operational controls for ML systems at scale.
  • Experience driving cross-team tooling or workflow standardization in ML or data platforms.
  • Experience working in a high-growth environment with multiple stakeholders and evolving ML platform needs.

What the JD emphasized

  • ML systems reliably at scale
  • ML lifecycle
  • production ML systems
  • model governance practices
  • ML tooling and workflows
  • ML systems
  • ML infrastructure
  • ML development
  • ML workflows
  • model training and batch inference
  • ML systems
  • ML platform
  • ML or data platforms
  • ML systems at scale
  • ML platform needs

Other signals

  • MLOps platform
  • ML lifecycle
  • training, orchestration, and deployment
  • governance practices
  • ML systems at scale