Aiml - Sr Machine Learning Engineer, Data and ML Innovation

Apple Apple · Big Tech · Cupertino, CA +1 · Machine Learning and AI

This role focuses on ensuring the quality and reliability of foundation models by designing, implementing, and maintaining evaluation infrastructure. The engineer will work closely with researchers to translate evaluation insights into actionable improvements and leverage agentic LLM systems for evaluations. The primary focus is on the evaluation gate (L5), with a secondary focus on post-training aspects (L2) as evaluation insights inform model improvements.

What you'd actually do

  1. Ensure the stability, reliability, and performance of Apple's foundation model evaluation system.
  2. Design and implement novel evaluation methodologies.
  3. Help design and implement tooling to simplify metrics generation, ingestion, and reporting.
  4. Leverage agentic LLM systems to facilitate and improve model evaluations.

Skills

Required

  • 5+ years of hands on ML engineering experiences
  • at least 1+ years working directly on large language models or generative AI
  • Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, or a related technical field — or equivalent practical experience
  • Strong software engineering fundamentals: debugging, testing, code reviews, and production reliability / scalability
  • Hands-on experience with LLM training and / or evaluation workflows

Nice to have

  • Hands on experience with evaluating large language models at scale or designing large language model benchmarks
  • Strong communication skills
  • Self-motivated and curious
  • High level of creative and critical thinking skills
  • High tolerance for ambiguity
  • ability to identify the most important problems to solve

What the JD emphasized

  • foundation model evaluation
  • foundation model performance
  • evaluation infrastructure
  • ML researchers
  • model hillclimbing
  • measuring model performance
  • foundation model lifecycle
  • LLM training and / or evaluation workflows
  • pre-training, post-training, online evaluation, offline evaluation, automated evaluation, human evaluation
  • evaluating large language models at scale
  • designing large language model benchmarks

Other signals

  • foundation model evaluation
  • evaluate foundation model performance
  • design and implement novel evaluation methodologies
  • ensure foundation model performance can be measured quickly and reliably