Postdoctoral Scholar, AI Evaluation & Standards

Johnson & Johnson Johnson & Johnson · Pharma · Madrid, Spain +5

Postdoctoral researcher to design methods for testing Generative AI tools in pharmaceutical R&D, focusing on accuracy, evidence-groundedness, traceability, usability, and appropriateness. The role involves developing evaluation frameworks, benchmark datasets, and analyzing LLM, RAG, and agent performance to ensure quality and safety before deployment.

What you'd actually do

  1. Design evaluation frameworks, rubrics, and criteria for GenAI tools used across pharmaceutical R&D.
  2. Develop therapeutic-area-specific criteria with business and scientific teams to reflect domain and use-case quality needs.
  3. Build benchmark datasets, reference answer sets, annotation guides, and evaluation datasets.
  4. Run expert reviews with scientific, clinical, regulatory, medical, data science, and engineering teams.
  5. Test LLM, RAG, and agent performance, including accuracy, source grounding, retrieval quality, citation fidelity, task completion, robustness, safety, and usability.

Skills

Required

  • GenAI evaluation
  • Python
  • data science
  • benchmarking
  • rubric design
  • biomedical research
  • technical communication

Nice to have

  • RAG evaluation
  • agent evaluation
  • biomedical informatics
  • expert review
  • annotation protocols
  • therapeutic area expertise
  • responsible AI
  • LLM APIs
  • embeddings
  • vector databases
  • prompt engineering
  • agent frameworks
  • AI evaluation tools
  • retrieval quality
  • generated answers
  • multi-step workflows
  • tool use
  • scientific reasoning
  • citation quality
  • evidence-grounded outputs
  • annotation instructions
  • adjudication processes
  • inter-rater reliability analyses
  • biomedical data standards
  • structured scientific or clinical data
  • ontologies
  • knowledge graphs
  • CDISC
  • FHIR
  • publications or applied research in AI evaluation, NLP, biomedical informatics, machine learning, data science, computational biology, bioinformatics, or a related field.

What the JD emphasized

  • PhD or equivalent research experience
  • Understanding of biomedical science, pharmaceutical R&D, therapeutic area science, translational science, clinical development, regulatory science, biomedical informatics, data science, or related areas.
  • Experience translating expert judgment into criteria, rubrics, datasets, protocols, or measurable outcomes.
  • Experience designing or applying evaluation methods, benchmark datasets, annotation protocols, validation studies, quality reviews, or assessment frameworks.
  • Interest in testing GenAI systems, including LLMs, RAG, and AI agents.
  • Proficiency in Python and common data science or machine learning tools.
  • Ability to analyze model outputs, compare performance, identify failure patterns, and recommend improvements.
  • Clear written and verbal communication skills.

Other signals

  • design evaluation frameworks for GenAI tools
  • test LLM, RAG, and agent performance
  • analyze failure patterns
  • translate evaluation findings into improvements