Senior Research Scientist | Model Scaling

DeepL DeepL · AI Frontier · London, United Kingdom · Research

Senior Research Scientist role focused on scaling foundation models for DeepL's next-generation translation systems. Responsibilities include selecting and evaluating open foundation models, leading architecture decisions for models with hundreds of billions of parameters, designing adaptation strategies using LoRA/PEFT, and owning the modeling lifecycle from prototyping to production delivery. Requires strong experience with fine-tuning and scaling LLMs, understanding architecture trade-offs, and Python/PyTorch/JAX/Tensorflow skills.

What you'd actually do

  1. Drive the selection and evaluation of open foundation / open-weight models as the basis for our next-generation translation systems.
  2. Lead model selection and general architecture decisions for scaling to hundreds of billions of parameters, including Mixture-of-Experts and other sparse or efficient designs.
  3. Design multi-capability adaptation strategies using LoRA, PEFT, and related methods.
  4. Own the modelling lifecycle for your work: prototyping, ablations, scaling experiments, evaluation, and delivery into production, with rigorous and reproducible evaluation.
  5. Partner closely with post-training, RL/RLHF, and instruction-following specialists to integrate alignment and capability work into the base model.

Skills

Required

  • adapting and scaling large language models via fine-tuning, instruction-tuning, or post-training of multi-billion-parameter models beyond black-box use
  • architecture trade-offs at scale (e.g. dense vs. MoE)
  • parameter-efficient and multi-capability adaptation (LoRA/PEFT and variants)
  • training models, running experiments, and debugging pipelines
  • carry research results through to production with engineering
  • Python
  • PyTorch/JAX/Tensorflow
  • communicate clearly
  • collaborate across teams
  • align research work with product and engineering priorities

Nice to have

  • quantifying uncertainty in large models — calibration and confidence estimation via Bayesian methods, ensembling, steering, or prompt-based approaches
  • machine translation
  • multilingual NLP
  • document-/layout-aware modelling experience
  • MoE-specific training and adaptation (e.g. expert routing, Mixture-of-LoRA-Experts)
  • large-scale data-mixture design

What the JD emphasized

  • scaling to hundreds of billions of parameters
  • LoRA, PEFT
  • delivery into production
  • adapting and scaling large language models via fine-tuning, instruction-tuning, or post-training of multi-billion-parameter models beyond black-box use
  • architecture trade-offs at scale

Other signals

  • scaling foundation models
  • architecture decisions
  • adaptation strategies
  • prototyping
  • large-scale experiments
  • delivery into production