Skip to content
Review

Enhancing Clinical Trial Analysis through Large Language Models for Multi-Evidence Natural Language Inference

2026 · International Conference on Language Resources and Evaluation · pp. 515-523 · 0 citations · 36 references
Computer Science

TL;DR

It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical decision-making.

View source

Similar papers

Open access 2026

Assessing the Difficulty of Inference Types in Natural Language Inference for Clinical Trials

Large Language Models (LLMs) achieve competitive results on Natural Language Inference when applied to clinical trials; however, it is not yet clear which type of inference LLMs perform well or poorly on. We address this by proposing new supplementary annotations for the existing NLI4CT dataset on the types of inferenc...

Mathilde Aguiar, Pierre Zweigenbaum, Nona Naderi · 0 citations
Review Open access Oct 2024

Large Language Model Benchmarks in Medical Tasks

With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets used in medical LLM tasks. These datasets span multiple modalities including t...

L. K. Yan, Qian Niu, Ming Li et al. · 32 citations · ⚡1
Review Open access Aug 2026

Evaluating large language model performance in US FDA regulatory science

This study compares the performance of three LLMs in extracting, analyzing, and synthesizing regulatory and clinical information from FDA drug reviews, guidance for the industry, and drug labels as accessed through their standard user interfaces, using antibiotics approved for complicated urinary tract infections betwe...

Khulud Bukhari, R. Rodriguez-Monguio, B. Lopez-Bermudez et al. · 0 citations
#artificial intelligence Review Sep 2026

A primer on evaluation methods for large language models in healthcare

This article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs by describing underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question.

S. E. McKinney, P. Vu, S. Justice et al. · 0 citations
Review Open access Sep 2026

Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews.

A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.

Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.