Skip to content
Review Open access

Benchmarking and developing large language models using one million clinical trials

Jul 2026 · npj Digital Medicine · 1 citation · ⚡ 1 influential

TL;DR

A large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature is introduced, demonstrating the potential of domain-adapted AI to improve evidence synthesis and clinical trial design and establishing as a foundation for scaling AI in clinical research.

Abstract

Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation for model benchmarking and development. Here, we introduce , a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing as a foundation for scaling AI in clinical research.

Read PDF

Similar papers

Review Open access Jul 2026

Tutorial: guidance on the use of large language models for medical research

This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.

Qiao Jin, Nicholas Wan, Robert Leaman et al. · 1 citation
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis, and multiple recommendations for future research are contained to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations
Preprint Aug 2026

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

Praveen Reddy, C. Mandke, Suvrankar Datta et al. · 0 citations
Review Open access Jul 2026

Artificial intelligence for genomic science: a scoping review of concepts, architectures, applications, and open challenges

This scoping review mapped how AI is defined and operationalized in genomic science, including machine learning, deep learning, graph-based methods, foundation models, and large language models, and synthesized their data modalities, applications, evaluation practices, interpretability strategies, and governance challenges.

W. Rodrigues, M. Parise, Doglas Parise et al. · 1 citation
Review Jul 2026

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

A scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration.

Zheng Tong, Yang Liu, Wanshu Fan et al. · 0 citations