Clinical Trials Benchmark Large Language Models
- •Researchers aggregate 1.6M clinical trial records from fifteen global registries for LLM development
- •Study builds 152K training and testing samples across eight clinical research tasks
- •An 8B domain-adapted LLM outperforms 70B generic counterparts across all eight tasks
Zifeng Wang, Jiacheng Lin, Qiao Jin and co-authors published a 2026 npj Digital Medicine study on benchmarking and developing large language models for clinical research using a structured resource built from 1.6M clinical trial records across fifteen global registries. The resource links trial records with biomedical ontologies and literature, giving researchers a data foundation for model benchmarking and development in clinical research.
The authors used the resource to build 152K training and testing samples across eight clinical research tasks, including systematic review, trial design and trial optimization. Benchmarking cutting-edge LLMs found limited clinical reasoning capability in generic LLMs, according to the study.
An 8B LLM developed with supervised fine-tuning and reinforcement learning outperformed 70B generic counterparts across all eight tasks. The reported relative improvements were 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, and the authors said the results show domain-adapted AI can improve evidence synthesis and clinical trial design.