Aiml Researcher - Foundation Model, Post-training

Apple Apple · Big Tech · Zurich, Zurich, Switzerland · Machine Learning and AI

Research role focused on post-training foundation models (LLMs) at Apple, aiming to transform them into capable assistants. Responsibilities include designing post-training strategies (RL, preference optimization, steering, safety), data generation (human/synthetic, curriculum learning), and developing robust evaluation methodologies. The role partners with pre-training and product teams. Requires expertise in deep learning, LLMs, post-training, or RL, with Python and JAX/PyTorch proficiency. Preferred qualifications include experience with large-scale distributed training, complex reasoning tasks, and transformer architectures.

What you'd actually do

  1. Design and iterate on end-to-end post-training strategies (including Reinforcement Learning) to unlock model capacities toward achieving specific model behaviours.
  2. Pioneer novel algorithms for preference optimization, model steering, and safety.
  3. Drive our data strategy by researching methods for high-quality human and synthetic data generation, automated data filtering, and curriculum learning to improve instruction following and reasoning.
  4. Design robust evaluation methodologies to measure model helpfulness, factuality, and utility, moving beyond static benchmarks to accurately capture real-world performance.
  5. Partner closely with pre-training teams to inform architecture choices, and with product teams to translate user requirements into model capabilities.

Skills

Required

  • deep learning
  • LLMs
  • post-training
  • reinforcement learning
  • Python
  • JAX or PyTorch
  • Masters/PhD or equivalent practical experience

Nice to have

  • large models at scale
  • distributed training
  • complex reasoning tasks
  • math
  • coding
  • logic
  • transformer architectures
  • communication skills
  • cross-functional collaboration

What the JD emphasized

  • post-training
  • reinforcement learning
  • instruction following
  • reasoning
  • evaluation methodologies
  • preference optimization
  • model steering
  • safety
  • human and synthetic data generation
  • curriculum learning

Other signals

  • foundation models
  • LLMs
  • post-training
  • instruction following
  • tool use
  • reasoning
  • evaluation