Sep 2026· Journal of Medical Internet Research· Vol 28· 0 citations· 31 references
Medicine
TL;DR
LLM agents perform reliably for question generation and SAP drafting but require expert verification of formula composition, cohort boundary logic, and concordance computation before results are reported.
Abstract
Abstract Background Large language model (LLM) agents capable of generating and executing statistical code from natural language may broaden access to clinical data analysis, yet which pipeline stages they perform reliably and which require expert oversight remain poorly defined. Objective This study aimed to evaluate the performance and systematic failure modes of an LLM agent across 5 stages of a clinical data analysis workflow. Methods The publicly available dataset and R script (R Foundation for Statistical Computing) were drawn from a previously published study of 12-year outcomes in 7802 patients with eyes with neovascular age-related macular degeneration at Moorfields Eye Hospital. Participants were evaluated using an LLM agent (Claude; Anthropic) across 3 interaction modes (Chat, Code, and Cowork). It was asked to perform 3 levels of data analysis practice: prompt A, to generate research questions from raw data only; prompt B, to develop a statistical analysis plan (SAP) from a high-level clinical objective, then execute it; and prompt C, to execute an analysis given an investigator-drafted SAP. Each was replicated 3 times (27 total runs). Qualitative evaluation of research question thematic coverage (prompt A), SAP completeness against a reference checklist (prompt B), and evaluation of execution outputs against validated reference values and of result text and narrative summaries against execution logs (prompts B and C) was conducted. Results The agent generated 18 clinically grounded questions spanning 7 domains; Cowork mode uniquely reached 3 thematic areas requiring data-driven methods. All 9 SAPs correctly identified the statistical framework. Kaplan-Meier estimates were near-identical across 17 completed runs. Systematic execution errors emerged: SAP quality did not predict code correctness, and within-mode errors propagated identically across independent repetitions. Result text accurately reflected execution logs in nearly all runs, though unit propagation and an undisclosed postcrash rerun were identified. Of 17 narrative summaries, 8 were fully satisfactory; 2 runs produced clinically meaningful errors. Conclusions LLM agents perform reliably for question generation and SAP drafting but require expert verification of formula composition, cohort boundary logic, and concordance computation before results are reported. Using an ophthalmology dataset as a controlled testbed, this study develops and applies an evaluation framework whose lessons are likely applicable across clinical specialties.
It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...
Shobanapriyan Chandrasegaran, Amal Htait· International Conference on...· 0 citations
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the execut...
Erfan D. Dehkalani, S. Shankaran, Abbot R. Laptook et al.· 0 citations
This article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs by describing underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question.
S. E. McKinney, P. Vu, S. Justice et al.· 0 citations
Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide...
Elena A. Mourelatou, Ioannis Katakis· bioRxiv· 0 citations
This work introduces a scalable, resource-efficient, and high-performance information extraction pipeline that leverages large language models (LLMs) to address challenges of free-text clinical records and develops a multi-dimensional assessment for deployment in data extraction tasks.
A. Y. Ong, Quang Nguyen, I. Barai et al.· npj Digital Medicine· 1 citation
Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provid...
Utkarsh Soni, S. Murtaza, Yi-Fan Nie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.