Senior Applied Scientist, Multilingual AI Evaluation

Apple Apple · Big Tech · Seattle, WA +2 · Machine Learning and AI

Senior Applied Scientist focused on developing and extending multilingual AI evaluation methodology and tooling for LLMs and agentic systems at Apple. The role involves designing, validating, and productionizing evaluation methods, benchmarks, and datasets across various languages and cultures, with a strong emphasis on linguistic expertise and Python implementation.

What you'd actually do

  1. Extend Apple’s AI evaluation methodology and tooling to new languages and locales, so our AI experiences are equally capable, accurate, and culturally appropriate — not just translated English.
  2. Design and validate evaluation methods, benchmarks, and metrics that capture language- and culture-specific phenomena: grammar, morphology, script, register, dialect, code-switching, and cultural norms.
  3. Build and curate high-quality multilingual datasets and human evaluation protocols, partnering with linguists and native-speaker annotators.
  4. Investigate how LLMs and agentic systems behave across languages, identifying systematic failure modes, capability gaps, and quality disparities between high- and low-resource languages.
  5. Partner with engineers to productionize your methods so they run reliably and at scale, implementing your own work in Python.

Skills

Required

  • MS in Linguistics, Computational Linguistics, NLP, Computer Science, or a related field — or equivalent research/work experience.
  • Deep expertise in linguistics, with working fluency in the structure of multiple languages beyond English.
  • Strong proficiency in Python.
  • Solid understanding of LLMs and AI evaluation fundamentals, including how language models process and generate across languages.
  • Demonstrated experience shipping or evaluating features across multiple languages or locales.
  • Experience designing benchmarks, datasets, or human evaluation protocols, with attention to statistical rigor and reproducibility.
  • Ability to drive initiatives independently and collaborate across a cross-functional, interdisciplinary team.
  • Strong written and verbal communication skills.

Nice to have

  • PhD in Linguistics, Computational Linguistics, or NLP with a focus on multilingual or cross-lingual modeling.
  • Publications in NLP, multilingual evaluation, or evaluation methodology.
  • Hands-on experience with modern ML frameworks (PyTorch, JAX) and with fine-tuning or evaluating LLMs.
  • Experience with low-resource languages, dialectal variation, or sociolinguistics.
  • Familiarity with localization/internationalization workflows and quality assessment.
  • Experience with LLM-as-judge approaches, rubric design, or bias and fairness evaluation across languages.
  • Fluency or professional proficiency in one or more languages in addition to English.

What the JD emphasized

  • evaluation methodology
  • multilingual
  • languages
  • cultures
  • evaluation tooling
  • language models
  • agentic systems
  • evaluation

Other signals

  • evaluation methodology
  • multilingual AI
  • LLMs
  • agentic systems