Senior Engineer I/ii, Drug Discovery Platform

Lila Sciences Lila Sciences · AI Frontier · Alewife, Cambridge, MA · Physical Sciences AI

Senior Engineer to own the data backbone for an AI-driven drug discovery platform. This includes building and operating the molecule queue, compound registration system, synthesis constraints, and experimental result capture pipelines that feed predictive models. The role focuses on ensuring low latency and high fidelity data for AI learning cycles.

What you'd actually do

  1. Design and operate the shared queue that brokers compounds between AI Scientists and the Make-Test platform, including the status state machine, batch grouping, priority semantics, and the read/write contracts both sides depend on.
  2. Build the canonical molecule registry, including SMILES/InChI normalization, stereochemistry handling, salt/parent resolution, duplicate detection, and stable internal IDs (e.g., LILA-NNN) that every downstream system can rely on.
  3. Build pipelines that ingest data from ChemSpeed automated synthesis, QC instruments (LCMS, NMR), and bioassay readers, capturing identity, purity, dose-response curves, IC50/EC50, ADMET measurements, and selectivity panels into a queryable store.
  4. Maintain the live state that the Batch Assembly AI reads each cycle, covering building-block inventory, advanced precursors, ChemSpeed capacity, stock alerts, and chemistry-specific constraints (reaction types, step budgets, solvent compatibility).
  5. Own data quality, lineage, schema evolution, and SLAs for the closed-loop cycle time target of a few days from computational proposal to experimental truth.

Skills

Required

  • Bachelor's or Master's in Computer Science, Chemistry, Computational Biology, or a related field
  • 5+ years building data platforms in production
  • Designed and shipped data platform components from the ground up
  • fluent in backend production APIs/services
  • Python and SQL
  • writes production-quality code
  • Production experience with relational and/or NoSQL databases
  • schema design for evolving scientific data
  • query optimization
  • comfortable across structured, semi-structured, and unstructured data
  • Comfortable modeling chemical and biological data
  • Experience with AWS
  • containerized deployment (Kubernetes)

Nice to have

  • some exposure to scientific or lab-generated data
  • Hands-on with RDKit, OpenEye, or equivalent
  • Direct experience building or operating a corporate compound registry
  • Familiarity with electronic lab notebooks, LIMS, and instrument data formats
  • Experience with Flyte, Airflow, Dagster, or Temporal
  • Lakehouse Patterns
  • Experience capturing data from automated synthesis platforms
  • Prior work in pharma, biotech, or academic drug discovery
  • Experience building data infrastructure that serves agentic and LLM-driven workflows

What the JD emphasized

  • data backbone
  • low latency and high fidelity
  • feed back into our predictive models
  • how fast our AI learns

Other signals

  • AI-driven drug discovery factory
  • data backbone for closed loop
  • feed back into our predictive models
  • how fast our AI learns