Senior / Staff Machine Learning Engineer

Apple Apple · Big Tech · Seattle, WA +2 · Software and Services

The Apple Cloud AI Platform team is seeking a Senior/Staff Machine Learning Engineer to build and scale the ML infrastructure and data platform that powers Apple's generative AI applications. This role involves designing and implementing large-scale data pipelines, SDKs, and distributed systems for data ingestion, versioning, governance, and efficient feeding of GPU fleets for model training and inference. The focus is on enabling ML engineers and researchers to build and ship models at scale, supporting foundation models, multimodal data, and RAG systems.

What you'd actually do

  1. Design and build the platform behind Apple's largest model builds — ingestion, immutable versioning, lineage, and governance across structured, unstructured, and multimodal data at petabyte scale, so every model run is reproducible from a versioned dataset
  2. Develop and evolve Python SDKs and core data libraries that ML engineers depend on to access, transform, and load model-ready datasets across every stage of model development
  3. Build high-throughput data access and loading primitives that feed Apple's largest GPU fleets, keeping workloads compute-bound rather than I/O-bound
  4. Build and operate distributed data pipelines spanning Spark, Daft, and Rust-based systems for ingestion, transformation, and large-scale data preparation
  5. Optimize platform components for tight integration with leading ML frameworks — PyTorch, JAX, and TensorFlow — so dataset access is a first-class concern in the model development loop

Skills

Required

  • machine learning
  • end-to-end ML workflow
  • data preparation
  • pipeline development
  • experimentation
  • evaluation
  • deployment
  • large scale distributed systems
  • modern generative techniques
  • transformers
  • diffusion
  • retrieval-augmented generation
  • data and machine learning infrastructure
  • fine-tuning workflows
  • model optimization
  • preparing models for scalable inference
  • generative AI
  • configuring, deploying and troubleshooting large scale production environments
  • designing, building, and maintaining scalable, highly available systems
  • Java
  • Python
  • Go
  • collaboration
  • communication
  • navigating ambiguity
  • evolving technical landscapes

Nice to have

  • PyTorch
  • JAX
  • TensorFlow
  • data loading
  • dataset access layer
  • Parquet
  • Iceberg
  • Delta
  • Lance
  • Ray Data
  • NVIDIA DALI
  • WebDataset
  • Mosaic StreamingDataset
  • performance engineering for I/O-bound workloads
  • Arrow
  • zero-copy
  • memory mapping
  • async I/O
  • high-throughput object storage access patterns at GPU scale
  • DataHub
  • OpenLineage
  • Unity Catalog
  • Spark
  • Daft
  • Polars
  • DuckDB internals
  • Docker
  • Kubernetes

What the JD emphasized

  • petabyte scale
  • large scale distributed systems
  • real-world production environments
  • large scale production environments
  • scalable, highly available systems

Other signals

  • ML infrastructure
  • generative AI
  • data pipelines
  • ML workflows
  • petabyte scale
  • Python SDKs
  • GPU fleets
  • Spark
  • Daft
  • Rust
  • PyTorch
  • JAX
  • TensorFlow
  • foundation models
  • multimodal data
  • retrieval-augmented systems