Senior Software Engineer - ML Infrastructure

Applied Intuition Applied Intuition · Robotics · Sunnyvale, CA · Machine Learning

Senior Software Engineer focused on ML Infrastructure, building and integrating end-to-end ML pipelines, distributed cloud GPU training, and evaluation systems. The role spans the entire ML lifecycle, working with modeling teams to solve complex data problems at scale and contribute to a company-wide platform for ML training, evaluation, and deployment.

What you'd actually do

  1. Design and implement distributed cloud GPU training approaches for deep learning model training and evaluation
  2. Build end-to-end machine learning pipelines and integrate them into core product workflows
  3. Encourage change, especially in support of ML engineering best practices, and maintain a high standard of excellence
  4. Collaborate with engineers across the entire company to solve complex data problems at scale

Skills

Required

  • 3+ years of professional experience
  • Experience with building software components to address production, full-stack machine learning challenges
  • Knowledge of the open source landscape with judgment on when to choose open source versus build in-house
  • Excellent analytical and problem-solving skills

Nice to have

  • Experience with developing, running, and managing orchestration systems like Airflow and Flyte that non engineers can use to build data pipelines
  • Experience with ML modeling frameworks (PyTorch, Tensorflow, etc.)
  • Experience with model serving platforms (TorchServe, TensorFlow Serving, NVIDIA Triton inference server, etc.)

What the JD emphasized

  • production, full-stack machine learning challenges
  • Opinions about building a company-wide platform for ML training, evaluation, and deployment

Other signals

  • ML infrastructure
  • ML pipelines
  • distributed cloud GPU training
  • ML lifecycle
  • production ML challenges