Scientist II / Senior ML Scientist, Data-efficient Learning for Drug Discovery

Lila Sciences Lila Sciences · AI Frontier · Alewife, Cambridge, MA +1 · Physical Sciences AI

Research Scientist role focused on developing data-efficient ML models for drug discovery, specifically in low-data regimes. The role involves building models, designing data acquisition strategies, and developing approaches like active learning, meta-learning, and uncertainty estimation. It connects ML model training with scientific decision-making and aims to build closed-loop learning workflows. The role also involves building multimodal models and translating predictions into practical recommendations.

What you'd actually do

  1. Build ML models that perform well in low-data regimes for drug discovery and molecular optimization.
  2. Design data acquisition strategies that identify which compounds, assays, DEL selections, simulations, structural predictions, or experiments should be run next to maximize learning.
  3. Develop active learning, meta-learning, fine-tuning, transfer learning, uncertainty-aware modeling approaches for focused chemical spaces.
  4. Train models on low-quantity, high-quality datasets generated by Lila's experimental, computational, and agentic discovery systems.
  5. Build multimodal models that can integrate DEL data, simulation outputs, assay data, protein and structural information, chemical features, literature or text-derived signals, images, and experimental metadata.

Skills

Required

  • PhD or equivalent experience in machine learning, computational chemistry, computational biology, statistics, computer science, bioengineering, or a related field.
  • Strong experience training ML models in low-data regimes.
  • Experience with active learning, Bayesian optimization, experimental design, meta-learning, fine-tuning, transfer learning, uncertainty estimation, or related data-efficient learning methods.
  • Experience building ML models for scientific, molecular, biological, chemical, pharmacological, biochemical, or other high-dimensional experimental datasets.
  • Experience with multimodal learning or methods that combine heterogeneous data sources.
  • Ability to reason about data acquisition strategy, not only model fitting.
  • Strong scientific judgment and ability to connect model behavior to experimental decisions.
  • Practical experience with PyTorch, JAX, scikit-learn, or equivalent ML tools.
  • Ability to collaborate across ML, data, computational science, experimental, and drug discovery teams.

Nice to have

  • Drug discovery experience, especially in molecular optimization, screening, or design-make-test-learn workflows.
  • General understanding of pharmacology, biochemistry, or mechanisms of molecular activity.
  • Experience with DEL, high-throughput screening, medicinal chemistry, assay data, simulation-derived features, protein or structure-based features, text or literature features, or scientific images.
  • Experience with closed-loop experimentation, autonomous labs, or agent-driven scientific workflows.
  • Experience with generative molecular design, candidate prioritization, or batch selection workflows.
  • Familiarity with causal inference, optimal experimental design, decision theory, or Bayesian methods.
  • Comfort working with frontier ML techniques where standard out-of-the-box approaches are insufficient.

What the JD emphasized

  • low-data regimes
  • data-efficient learning
  • active learning
  • meta-learning
  • uncertainty estimation
  • experimental design
  • multimodal modeling
  • closed-loop systems
  • agentic discovery systems
  • low-quantity but high-quality datasets
  • data acquisition strategy

Other signals

  • low-data regimes
  • active learning
  • meta-learning
  • uncertainty estimation
  • experimental design
  • multimodal modeling
  • closed-loop systems