Senior Research Scientist | Model Steering

DeepL DeepL · AI Frontier · London, United Kingdom · Research

Senior Research Scientist to lead fine-tuning, post-training, model-steerability, and reinforcement learning for DeepL's next generation of LLM-based translation models. This role involves driving research, prototyping, running large-scale experiments, and integrating models into production, with a focus on blending data, shaping model behavior, and making models steerable based on user preferences and context. The role also includes building reward models, handling multimodal content, and establishing evaluation and monitoring practices.

What you'd actually do

  1. Drive the development of translation models that are steerable conditioned on user preferences, rules and context.
  2. Drive hands-on research and development on post-training for our core translation models: supervised fine-tuning, knowledge distillation, preference optimization, and reinforcement learning tuned to translation quality.
  3. Build reward models and evaluator models for translation, including rubric- and reference-based grading, and investigate and mitigate reward hacking and quality-estimation failure modes.
  4. Drive an agenda toward models that ingest multimodal content and context to increase translation quality.
  5. Own the full lifecycle of model delivery: prototyping, ablations, training, evaluation, optimization, and production deployment, working closely with engineering to ship into real-time systems at scale.

Skills

Required

  • LLM post-training (SFT, DPO)
  • knowledge distillation
  • reinforcement learning (RLHF/RLAIF, PPO/GSPO)
  • reward modeling
  • synthetic data pipelines
  • preference data pipelines
  • LLM-as-judge evaluation
  • data curation
  • data filtering
  • Python
  • PyTorch
  • JAX
  • Tensorflow

Nice to have

  • distributed/multi-node training
  • FSDP
  • DeepSpeed
  • Megatron-style frameworks
  • efficient training techniques
  • multimodal content ingestion

What the JD emphasized

  • Proven experience making large models steerable and instruction-following by identifying the most effective method to instill a given behavior, drawing from instruction tuning, latent space methods, steering vectors, and/or constrained encoding and decoding methods.
  • Deep, hands-on expertise in LLM post-training (SFT, DPO), knowledge distillation (teacher-student training), and/or reinforcement learning (RLHF/RLAIF, PPO/GSPO, and reward modeling).
  • Strong data-centric instincts for building synthetic-data and preference-data pipelines, LLM-as-judge generation, data curation and filtering, and reasoning about data mixtures and ablations.
  • Experience designing evaluation and reward signals using automatic metrics, LLM-as-judge evaluation, non-verifiable rewards, and human-in-the-loop evaluation.
  • A hands-on builder who enjoys training models, running experiments, debugging pipelines, and integrating ML systems into production while staying grounded in product impact and real-world quality.
  • Ownership of a substantial research direction with strong execution, and experience mentoring others on a fast-moving, applied research team.
  • Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow), and the ability to communicate clearly and align research with product and engineering priorities.

Other signals

  • LLM-based translation models
  • fine-tuning
  • post-training
  • model-steerability
  • reinforcement learning
  • large-scale experiments
  • production deployment