ML Scientist I/ii, Nucleic Acid Design

Lila Sciences Lila Sciences · AI Frontier · San Francisco, CA · AI

ML Scientist focused on designing RNA and DNA sequences using ML models, spanning regulatory and coding sequence contexts. The role involves developing de novo generation, sequence property prediction, diverse set selection, and active learning strategies, while investigating biological mechanisms and integrating models into an autonomous science platform. Collaboration with experimental scientists, ML researchers, and platform teams is key.

What you'd actually do

  1. Build ML models for RNA and DNA sequence design across regulatory and coding sequence contexts.
  2. Develop methods spanning de novo generation, sequence property prediction, diverse set selection for experimental validation, and active learning strategies.
  3. Deeply investigate the biological mechanisms of designed sequences and propose hypotheses about why they succeed or fail. Turn these insights into better models and future design principles.
  4. Partner with experimental scientists to propose informative assays, validation strategies, and learning loops.
  5. Collaborate with ML scientists and engineers across Lila to integrate nucleic acid design models into robust platforms and agent-driven frameworks.

Skills

Required

  • PhD or equivalent experience in machine learning, computational biology, bioengineering, computer science, statistics, or a related quantitative field.
  • Hands-on experience building, training, and evaluating ML models for DNA or RNA.
  • Strong foundation in modern ML methods, with practical experience using frameworks such as PyTorch, JAX, or equivalent tools.
  • Experience developing models for sequence design, sequence-function prediction, generative modeling, or active learning.
  • Ability to reason about complex biological systems, scope ambiguous scientific problems, and formulate ML approaches that address difficult sequence-function challenges.
  • Curiosity about nucleic acid biology, including RNA biology, regulatory genomics, or related sequence-to-function problems.
  • Strong communication and collaboration skills, with a preference for team-based science and the ability to build shared technical direction across ML, engineering, platform, and experimental teams.

Nice to have

  • Experience with regulatory element design, sequence-to-expression DNA models, or models trained on genomic, MPRA, STARR-seq, or related functional genomics data.
  • Experience with RNA sequence-function modeling, RNA secondary structure modeling, UTR design, or inverse design methods for RNA sequences.
  • Experience with sequence design in applied therapeutic contexts.
  • Experience collaborating with wet-lab teams to close the design-test-learn loop, including assay design, experimental prioritization, and interpretation of validation data.
  • Familiarity with high-throughput experimental datasets, pooled screens, reporter assays, or other sequence-function measurements.
  • Industry experience translating ML research into practical biological design workflows, experimental campaigns, or platform capabilities.

What the JD emphasized

  • Hands-on experience building, training, and evaluating ML models for DNA or RNA.
  • Ability to reason about complex biological systems, scope ambiguous scientific problems, and formulate ML approaches that address difficult sequence-function challenges.
  • Curiosity about nucleic acid biology, including RNA biology, regulatory genomics, or related sequence-to-function problems.

Other signals

  • ML models for RNA and DNA sequence design
  • sequence-function modeling
  • generative modeling
  • active learning strategies
  • autonomous science platform