Principal Scientist – Data Insights & Hypothesis Design

Tempus AI Tempus AI · Vertical AI · Chicago, IL +1 · Remote

Principal Scientist role focused on data insights and hypothesis design within a healthcare AI company. The role involves ensuring scientific coherence of a multimodal data platform, bridging molecular science and data engineering, and expanding applications into new therapeutic areas. Responsibilities include data evaluation, prototyping, client alignment, scientific advisory, and cross-domain data oversight, with a focus on scientific excellence and data integrity. Requires expertise in quantitative disciplines, precision medicine, data engineering, emerging AI/foundational models, multi-omics, and translational impact.

What you'd actually do

  1. Conduct rigorous data evaluations, write technical code (SQL, R, Python), and prototype new data offerings within the Tempus Data Model (TDM). Deliver implementation-ready table and column specifications to seamlessly transition initial concepts into scalable, production-grade offerings.
  2. Translate complex client research objectives across the drug discovery and development lifecycle into actionable clinico-genomic data strategies. Ensure seamless alignment between client hypotheses and Tempus capabilities, effectively matching the right datasets to their specific research needs and guiding how to best leverage them.
  3. Serve as a core scientific authority across the organization. Evaluate datasets, perform User Acceptance Testing (UAT), and build data insight reports and capability demonstrations that empower various teams with scientifically validated evidence.
  4. Support the continuous growth and scaling of Tempus' expanding clinical-genomic footprint across diverse disease domains, with an active focus on neurology and Immunology. Bring a holistic, cross-disciplinary view of human biology to profile and validate terabyte-scale datasets, leveraging a proven track record of scientific agility to seamlessly adapt, uncover novel biological patterns, and extract meaningful findings across varying disease areas.

Skills

Required

  • SQL
  • R
  • Python
  • PhD in a quantitative, NGS-related discipline (e.g., Computational Biology, Genetics, Bioinformatics, Systems Biology)
  • 6+ years of industry experience across precision medicine, translational research, scientific product development, or related commercial and operational areas.
  • Deep expertise in at least one of oncology, neurology, or immunology.
  • Strong, demonstrable experience analyzing large-scale clinico-genomic datasets across multiple disease areas.
  • Proven experience collaborating directly with data engineering and platform teams to build large-scale relational databases, design advanced SQL data models, and support the robust infrastructure needed to create and deploy massive, production-grade models.
  • Active, up-to-date expertise with cutting-edge computational modeling, including experience leveraging foundation models, generative AI architectures, and agentic workflows to interpret and navigate complex biological data.
  • Deep expertise executing multi-omic analyses across a diverse spectrum of high-throughput sequencing applications and biotechnologies, including specialized NGS modalities such as single-cell sequencing, immunomics, metagenomics, and varied sequencing platforms.
  • Proven experience bringing data to market and translating scientific analyses into production-grade data products or clinical-facing solutions.
  • Proficiency in Python and/or R for reproducible scientific analysis, workflow development, and data visualization.
  • Strong cross-functional communication skills with the ability to architect end-to-end execution plans, manage client expectations, and mentor technical peers through structured delegation.

What the JD emphasized

  • strong computational, statistical, and translational research skills
  • deep interest in reproducible science and artificial intelligence
  • hands-on technical execution
  • strategic delegation
  • scientific excellence, enablement, and data integrity
  • rigorous data evaluations
  • scalable, production-grade offerings
  • complex client research objectives
  • clinico-genomic data strategies
  • scientifically validated evidence
  • continuous growth and scaling
  • terabyte-scale datasets
  • scientific agility
  • novel biological patterns
  • quantitative, NGS-related discipline
  • precision medicine
  • translational research
  • large-scale relational databases
  • foundation models
  • generative AI architectures
  • agentic workflows
  • multi-omic analyses
  • high-throughput sequencing applications
  • bringing data to market
  • production-grade data products
  • clinical-facing solutions
  • reproducible scientific analysis
  • architect end-to-end execution plans

Other signals

  • advancing the healthcare industry
  • AI to impact clinical care
  • proprietary platform connects an entire ecosystem of real-world evidence
  • deliver real-time, actionable insights
  • bridging the gap between complex molecular science and scalable data engineering
  • expanding our multi-modal applications
  • large-scale datasets
  • hands-on technical execution
  • strategic delegation
  • scientific excellence, enablement, and data integrity
  • prototype new data offerings within the Tempus Data Model
  • scalable, production-grade offerings
  • Translate complex client research objectives
  • clinico-genomic data strategies
  • scientifically validated evidence
  • continuous growth and scaling
  • terabyte-scale datasets
  • novel biological patterns
  • quantitative, NGS-related discipline
  • precision medicine
  • translational research
  • large-scale relational databases
  • foundation models
  • generative AI architectures
  • agentic workflows
  • multi-omic analyses
  • high-throughput sequencing applications
  • bringing data to market
  • production-grade data products
  • clinical-facing solutions
  • reproducible scientific analysis
  • architect end-to-end execution plans