Senior Machine Learning Systems Engineer

Reddit Reddit · Consumer · United States · Remote · Machine Learning

Senior Machine Learning Systems Engineer to lead development of a platform for large scale ML models at Reddit, focusing on MLOps, graph ML codebase, performance tuning, and optimizing batch data processing. Requires 5+ years of experience in ML infrastructure, model training/deployment, and cloud technologies.

What you'd actually do

  1. Design end-to-end model lifecycle patterns (MLOps) to boost velocity of development for ML engineers, including data preparation, model management, experiment tracking, and more
  2. Zero-to-one development and support of a graph ML codebase and platform that abstracts away common patterns and enables greater model scalability and iteration
  3. Collaborate with ML engineers on performance tuning, including improving model training time, efficiency, and GPU training costs in a large, distributed ML training environment
  4. Optimize batch data processing within a data warehouse and with tools such as Apache Beam, Apache Spark, Ray Data, and more
  5. Architect pipelines to build and maintain massive graph data structures on the order of billions of nodes and tens of billions of edges

Skills

Required

  • Python
  • PyTorch
  • Tensorflow
  • GCP BigQuery
  • Google Cloud Storage
  • Terraform
  • MLflow
  • Wandb
  • Ray
  • Kubernetes
  • Apache Beam
  • Apache Spark
  • Ray Data

Nice to have

  • graph databases (Neo4j, JanusGraph, TigerGraph)
  • graph neural networks (GNNs)
  • PyTorch Geometric
  • Deep Graph Library

What the JD emphasized

  • 5+ years of experience in ML infrastructure
  • model training
  • model deployments
  • ML optimization
  • GPU profiling
  • cloud-based technologies for supporting an ML platform
  • MLOps tools
  • distributed training frameworks

Other signals

  • MLOps platform
  • large scale ML models
  • model lifecycle
  • distributed training
  • GPU training costs