Staff Research Engineer, Scientific Computing and Ml/physics Infrastructure

Lila Sciences Lila Sciences · AI Frontier · Alewife, Cambridge, MA +1 · Physical Sciences AI

Research Engineer role focused on transforming research tools and prototypes into scalable, efficient, and maintainable systems for drug discovery. This involves building and supporting ML and physics infrastructure, ensuring reliable execution across compute environments, optimizing GPU utilization, and architecting larger-scale systems around research code. The role emphasizes collaboration with scientists to create agent-usable scientific pipelines and packaging tools into reusable services or APIs.

What you'd actually do

  1. Take research tools, prototypes, and scientific workflows developed by scientists or academic-style researchers and make them scalable, efficient, and maintainable.
  2. Collaborate directly with computational biophysics, computational chemistry, and machine learning scientists to turn research workflows into scalable agent-usable systems.
  3. Build and support ML and physics infrastructure for model training, molecular simulation, data processing, and agent-executed scientific workflows.
  4. Ensure workflows run reliably across multiple clusters and compute environments.
  5. Improve GPU utilization, distributed execution, throughput, fault tolerance, and reproducibility for ML and scientific workloads.

Skills

Required

  • Strong software engineering skills in Python
  • experience working with ML, scientific computing, or simulation codebases
  • Experience building, scaling, or operating distributed systems for research, ML, physics, simulation, or data-intensive workloads
  • Practical knowledge of GPU computing, performance profiling, distributed execution, and failure modes in large-scale workloads
  • Experience with PyTorch, JAX, CUDA-aware workflows, or related ML/scientific computing frameworks
  • Practical knowledge of Linux, Docker or containers, dependency management, and reproducible development environments
  • Experience with orchestration, scheduling, or distributed execution systems such as Kubernetes, Slurm, Ray, Flyte, Argo, or similar tools
  • Ability to take prototype-quality research code and improve its architecture, scalability, reliability, and maintainability
  • Strong debugging skills across code, environments, infrastructure, data pipelines, and compute clusters
  • Ability to work directly with researchers, understand ambiguous technical needs, and convert them into robust engineering solutions

Nice to have

  • Familiarity with chemistry, computational biophysics, molecular simulation, computational chemistry, cheminformatics, or drug discovery workflows
  • Experience with cloud GPU infrastructure, multi-cluster execution, or hybrid compute environments
  • Experience building tools for LLM agents or automated research workflows
  • Experience with workflow observability, checkpointing, retries, and fault-tolerant scientific workloads
  • Experience with CI, testing, packaging, and release practices for research software
  • Comfort supporting fast-moving research teams without over-engineering exploratory work

What the JD emphasized

  • scalable agent-usable systems
  • agent-executed scientific workflows
  • agent-usable scientific pipelines

Other signals

  • ML infrastructure for drug discovery
  • agent-usable scientific pipelines
  • improve code quality, architecture, GPU efficiency, cluster portability, and operational reliability