Skip to content
Open access

Performance, Failures, and Oversight of a Large Language Model Agent for Clinical Data Analysis: Evaluation Study

Sep 2026 · Journal of Medical Internet Research · Vol 28 · 0 citations · 31 references
Medicine

TL;DR

LLM agents perform reliably for question generation and SAP drafting but require expert verification of formula composition, cohort boundary logic, and concordance computation before results are reported.

Abstract

Abstract Background Large language model (LLM) agents capable of generating and executing statistical code from natural language may broaden access to clinical data analysis, yet which pipeline stages they perform reliably and which require expert oversight remain poorly defined. Objective This study aimed to evaluate the performance and systematic failure modes of an LLM agent across 5 stages of a clinical data analysis workflow. Methods The publicly available dataset and R script (R Foundation for Statistical Computing) were drawn from a previously published study of 12-year outcomes in 7802 patients with eyes with neovascular age-related macular degeneration at Moorfields Eye Hospital. Participants were evaluated using an LLM agent (Claude; Anthropic) across 3 interaction modes (Chat, Code, and Cowork). It was asked to perform 3 levels of data analysis practice: prompt A, to generate research questions from raw data only; prompt B, to develop a statistical analysis plan (SAP) from a high-level clinical objective, then execute it; and prompt C, to execute an analysis given an investigator-drafted SAP. Each was replicated 3 times (27 total runs). Qualitative evaluation of research question thematic coverage (prompt A), SAP completeness against a reference checklist (prompt B), and evaluation of execution outputs against validated reference values and of result text and narrative summaries against execution logs (prompts B and C) was conducted. Results The agent generated 18 clinically grounded questions spanning 7 domains; Cowork mode uniquely reached 3 thematic areas requiring data-driven methods. All 9 SAPs correctly identified the statistical framework. Kaplan-Meier estimates were near-identical across 17 completed runs. Systematic execution errors emerged: SAP quality did not predict code correctness, and within-mode errors propagated identically across independent repetitions. Result text accurately reflected execution logs in nearly all runs, though unit propagation and an undisclosed postcrash rerun were identified. Of 17 narrative summaries, 8 were fully satisfactory; 2 runs produced clinically meaningful errors. Conclusions LLM agents perform reliably for question generation and SAP drafting but require expert verification of formula composition, cohort boundary logic, and concordance computation before results are reported. Using an ophthalmology dataset as a controlled testbed, this study develops and applies an evaluation framework whose lessons are likely applicable across clinical specialties.

Read PDF

Similar papers

Review 2026

Enhancing Clinical Trial Analysis through Large Language Models for Multi-Evidence Natural Language Inference

It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...

Shobanapriyan Chandrasegaran, Amal Htait · 0 citations
#artificial intelligence Preprint Sep 2026

Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation

Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the execut...

Erfan D. Dehkalani, S. Shankaran, Abbot R. Laptook et al. · 0 citations
#artificial intelligence Review Sep 2026

A primer on evaluation methods for large language models in healthcare

This article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs by describing underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question.

S. E. McKinney, P. Vu, S. Justice et al. · 0 citations
Review Open access Nov 2025

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide...

Elena A. Mourelatou, Ioannis Katakis · 0 citations
Open access Aug 2026

Developing a scalable pipeline for data extraction from clinical letters through resource-efficient prompt engineering

This work introduces a scalable, resource-efficient, and high-performance information extraction pipeline that leverages large language models (LLMs) to address challenges of free-text clinical records and develops a multi-dimensional assessment for deployment in data extraction tasks.

A. Y. Ong, Quang Nguyen, I. Barai et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provid...

Utkarsh Soni, S. Murtaza, Yi-Fan Nie et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.