Senior Software Engineer

Apple Apple · Big Tech · Cupertino, CA +1 · Machine Learning and AI

Senior Software Engineer for Siri organization focused on building evaluation systems, tooling, and infrastructure for ML models, reward modeling, and alignment signals. The role emphasizes an 'evals first' approach, requiring strong software engineering fundamentals with ML/LLM depth, and involves production debugging, observability, and reproducibility.

What you'd actually do

  1. Supporting the evaluation of new Siri features and interaction modalities, working from ambiguous early requirements toward concrete, automated coverage
  2. Turning product goals into measurable system behavior — instrumenting the product, building eval harnesses, and creating test datasets grounded in real user workflows
  3. Building and improving reward models and alignment signals that measure whether Siri responses meet user needs
  4. Designing and shipping evaluation tooling, pipelines, and architecture end-to-end — from data ingestion through scoring to monitoring in production — with an eye toward observability, logging, and reproducibility
  5. Diagnosing failures across the stack, from environment provisioning through pipeline execution to scoring — enabling auto-diagnostics and driving durable fixes by partnering across engineering, infrastructure, and program teams to align on interfaces, priorities, and shared standards

Skills

Required

  • Strong programming skills in one or more compiled languages (Swift, C++, or Objective-C)
  • Strong Python skills
  • Solid computer science fundamentals, including data structures, algorithms, and clean, testable code
  • Experience with backend/API development and production debugging
  • Excellent communication and cross-team collaboration skills, with experience working effectively within large, cross-functional organizations
  • M.S. or B.S. in Computer Science, Machine Learning, or a related field (or equivalent experience)

Nice to have

  • Experience evaluating ML, LLM, or agent-based systems, including familiarity with metrics, scoring methodology, trajectory and outcome analysis, and techniques like prompting, RAG, or LLM as judge
  • Understanding of reinforcement learning and the underlying techniques behind modern LLMs (e.g. transformer architectures, RLHF/RLAIF, fine-tuning, reward modeling ) and frameworks such as PyTorch or Hugging Face, as applied to evaluation and reward signal design
  • Familiarity with eval-driven development — defining success criteria and test cases from product goals and real user workflows rather than abstract benchmarks
  • Experience with data science methods applied to quality measurement — defining ground truth, measuring inter-rater agreement (e.g. Cohen's/Fleiss' kappa), and validating automated scorers using basic statistical techniques (e.g. correlation, confidence intervals, hypothesis testing)
  • Experience with MLOps, deployment, and test/eval environment management — containerization, CI/CD, model versioning, monitoring, cloud platforms (AWS, GCP, or similar), and staging or provisioning environments to produce repeatable, deterministic conditions
  • Ability to quickly learn and adapt to evolving technologies and tools, such as GenAI-assisted coding, new ML frameworks, and emerging LLM/agent tooling
  • Comfort communicating and collaborating effectively across multicultural teams and time zones

What the JD emphasized

  • evaluation quality
  • reward/alignment signals
  • data science rigor
  • evals first thinking
  • observability
  • reproducibility
  • largely unexplored space
  • self-driven
  • define your own path
  • moves fluidly across these areas
  • strong software engineering fundamentals
  • ML/LLM depth
  • AI-facing tooling
  • ambiguous early requirements
  • measurable system behavior
  • real user workflows
  • production
  • diagnosing failures
  • durable fixes
  • compiled languages
  • Python skills
  • computer science fundamentals
  • clean, testable code
  • quickly learn and adapt
  • GenAI-assisted coding
  • emerging LLM/agent tooling
  • backend/API development
  • production debugging
  • cross-team collaboration
  • large, cross-functional organizations
  • evaluating ML, LLM, or agent-based systems
  • metrics
  • scoring methodology
  • trajectory and outcome analysis
  • prompting
  • RAG
  • LLM as judge
  • reinforcement learning
  • modern LLMs
  • transformer architectures
  • RLHF/RLAIF
  • fine-tuning
  • reward modeling
  • PyTorch
  • Hugging Face
  • evaluation and reward signal design
  • eval-driven development
  • success criteria
  • test cases
  • product goals
  • real user workflows
  • abstract benchmarks
  • data science methods
  • quality measurement
  • ground truth
  • inter-rater agreement
  • automated scorers
  • statistical techniques
  • MLOps
  • deployment
  • test/eval environment management
  • containerization
  • CI/CD
  • model versioning
  • monitoring
  • cloud platforms
  • staging or provisioning environments
  • repeatable, deterministic conditions
  • communicating and collaborating effectively
  • multicultural teams and time zones

Other signals

  • building evaluation systems
  • ML engineering techniques
  • reward modeling
  • alignment signals
  • production debugging
  • observability
  • reproducibility