Aug 2026· Radiology· Vol 320 2, pp.
e253238
· 1 citation· 15 references
Medicine
TL;DR
LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation.
Abstract Background Although large language models (LLMs) have demonstrated the ability to generate the impression section from radiology findings automatically, the incremental diagnostic value of clinical information for these models remains unclear. Objective This study aimed to evaluate the incremental diagnostic value of clinical information for LLMs and compare their performance with that of radiologists. Methods This retrospective study included radiology reports from patients with histopathologically confirmed liver, lung, and breast diseases from 2 institutions between October 2021 and February 2025. We defined three progressive information input scenarios: (1) basic patient information and imaging findings, (2) scenario A plus chief complaint or clinical history, and (3) scenario B plus key laboratory results. Scenario-based data were input into 3 general-purpose LLMs (DeepSeek-R1, Gemini 2.5 Pro, and GPT-4o), generating 2709 entries. Diagnostic accuracy was assessed for both benign-malignant differentiation and disease diagnosis, with histopathology serving as the reference standard. Accuracy was compared among scenarios and against radiologist performance using the McNemar test, and P values were adjusted using the Holm-Bonferroni correction for multiple comparisons. Results A total of 301 patients with pathologically confirmed diseases were included (mean age 53.5, SD 12.0 years; women: n=208, 69.1%). In the liver cohort, a numerical trend toward higher accuracy was observed in scenario C compared with scenario A across all 3 models (scenario C range: 72.3%‐76.2% vs scenario A range: 64.4%‐68.3%); these differences did not reach statistical significance after Holm-Bonferroni correction (all adjusted P>.99). Notably, the DeepSeek-R1 model in scenario C achieved the highest diagnostic accuracy (77/101, 76.2%), with no evidence of a difference compared with radiologists (82/101, 81.2%; adjusted P>.99). In contrast, results in the lung and breast cohorts were more heterogeneous. In the lung cohort, GPT-4o achieved its highest accuracy in scenario A for disease diagnosis (68/92, 73.9%), which exceeded its performance in scenario B (64/92, 69.6%) and scenario C (66/92, 71.7%), suggesting that additional clinical information did not confer a consistent benefit. Gemini 2.5 Pro in scenario B achieved the highest accuracy in this cohort (72/92, 78.3%); however, no statistically significant difference was found compared with radiologists (80/92, 87.0%; adjusted P=.25). In the breast cohort, DeepSeek-R1 achieved the numerically highest diagnostic accuracy in scenario A, and it decreased numerically with the addition of laboratory tests, although no significant difference was found between scenarios A and C (73/108, 67.6% vs 71/108, 65.7%; adjusted P>.99). Conclusions While the addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.
Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al.· Journal of Medical Internet...· 0 citations
BACKGROUND
Chest radiography remains the most widely used imaging modality worldwide; however, its interpretation is inherently challenging because of overlapping anatomical structures and subtle findings. Recent advances in multimodal large language models (LLMs) have enabled automated radiology report generation, yet their clinical performance relative to domain-specific medical AI systems remains insufficiently validated.
OBJECTIVES
This study aimed to evaluate the performance and clinical applicability of a domain-specific multimodal AI model (M4CXR) compared with a general-purpose LLM (ChatGPT-4o) for chest radiograph interpretation.
METHODS
In this retrospective study, 500 anonymized chest radiographs from a single tertiary care center were analyzed. Four board-certified radiologists independently evaluated AI-generated reports from both models. Key outcomes included key finding detection (categorized as complete, partial, or inconsistent), report generation time, and report discrepancies assessed using the RADPEER scoring system. Agreement between original and M4CXR-assisted RADPEER scores was assessed using intraclass correlation coefficients and weighted Cohen's kappa. Statistical analyses included paired t-tests, and chi-square tests.
RESULTS
M4CXR demonstrated significantly higher report consistency than GPT-4o, with complete concordance observed in 55.8% versus 19.8% of cases, and lower inconsistency rates (25.2% vs. 46.4%, P<.001). The use of M4CXR significantly reduced report generation time compared with unaided interpretation (16.3 ± 12.9 s vs. 179.2 ± 50.4 s, P<.001). RADPEER-based discrepancy analysis revealed no significant differences between original and AI-assisted interpretations. Agreement between original and M4CXR-assisted RADPEER scores showed good reliability (ICC = 0.701), and weighted kappa analysis showed substantial agreement (κw = 0.652).
CONCLUSIONS
Domain-specific multimodal AI model evaluated in this study demonstrated higher diagnostic consistency than the general-purpose LLM evaluated under the study conditions. These findings suggest the potential of specialized AI models as viable assistive tools, while highlighting the complementary utility of general-purpose LLMs in broader clinical contexts. Future integration should prioritize human-AI collaboration and prospective multi-center validation.
Tae-Hoon Kim, J. Hong, Jihun Hyun et al.· BMC Medical Imaging· 0 citations
BACKGROUND
Large language models (LLMs) show promise for converting complex radiology reports into patient-centric language, but inherent output instability may limit clinical application.
OBJECTIVES
To quantitatively assess the translational accuracy, error rates, and instability of various LLMs when generating patient-centric radiology reports, and evaluate demographic influences on report readability.
MATERIALS AND METHODS
This retrospective study evaluated 320 de-identified radiology reports processed by three LLMs using a two-stage (baseline and optimized) prompt engineering strategy. Two senior radiologists evaluated medical accuracy, completeness, and recommendation suitability. Readability was evaluated by 16 non-medical participants stratified by age and education.
RESULTS
Professional radiological evaluation revealed that all tested models exhibited inherent instability, omitted information, and tended to generate risk-averse, generalized clinical recommendations. To address these limitations, optimized structured prompts significantly reduced model output variance and improved translational accuracy, with particularly prominent effects observed in DeepSeek-R1 and ChatGPT-4.0. Overall, large language models significantly enhanced the readability of radiology reports (P < 0.05), with DeepSeek-R1 achieving the best performance. However, patients' self-reported comprehension of the reports was affected by demographic characteristics.
CONCLUSION
Large language models can effectively improve the readability of radiology reports, yet all such models inherently suffer from output instability and information omission. Optimized structured prompting can substantially reduce the variability of model outputs and improve the accuracy of medical text translation. Nevertheless, LLMs should currently be strictly confined to human-supervised auxiliary tools rather than applied as standalone clinical solutions.
Yun Mao, Chunyan Wang, Wei Wang et al.· Academic Radiology· 0 citations
A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.
Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al.· Indian Journal of Computer S...· 0 citations
BACKGROUND
Reviewing pathology, imaging, and consultation documents in oncology can be time-consuming, particularly when records originate from external facilities in different file formats. This study aimed to evaluate the impact of a Retrieval-Augmented Generation (RAG)-enabled GPT-4o summarization agent on clinical workflows and quality of outside-record summaries in breast surgical oncology.
METHODS
Initial performance evaluation of a GPT-4o/RAG agent to generate summaries of oncologic reports in 50 charts followed by a prospective pilot test of sequential cases, with each AI summary evaluated using a modified Provider Documentation Summarization Quality Instrument (PDSQI-9; 1-5 Likert scale), including dichotomized ratings (low [1-3], high [4, 5]), binomial testing, frequency and type of user-reported errors, clinician-coded error criticality (treatment-impacting vs noncritical). Pre- and post-use survey of documentation burden (NASA TLX) and user experience was performed.
RESULTS
Among 62 cases, AI-generated summaries were rated high for accuracy, usefulness, succinctness, and source citation. Thoroughness without omission was rated low in 28 (45%) summaries. Errors were noted in 25 (40%) surveys, with 13 (52%) classified as critical (treatment-impacting). The most common error type involved imaging, reported in 17 (68%) cases. For perceived time savings, the median response was neutral, but qualitative feedback described the tool as helpful for straightforward cases and as reducing typing burden but requiring workflow adjustment and improvements for complex cases.
CONCLUSIONS
Although users rated RAG-enabled GPT-4o agent-generated documentation summaries favorably on several quality domains, they frequently lacked thoroughness and occasionally contained treatment-relevant errors. Human review and further iteration of the technology remain necessary before implementation.
Ko Un Park, Bergen K. Sather, A. Shah et al.· Annals of Surgical Oncology· 0 citations
This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.
Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al.· 0 citations