Aug 2026· Korean Journal of Radiology· Vol 27, pp. 829· 0 citations· 32 references
Medicine
TL;DR
Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements.
Abstract
Objective
To evaluate adherence to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) in radiology and medical imaging studies involving large language models (LLMs).
Materials And Methods
We conducted a cross-sectional audit of original LLM research studies published between January 1 and December 26, 2025, in Q1 journals within the Web of Science "Radiology, Nuclear Medicine, and Medical Imaging" category. PubMed and Scopus were searched to identify eligible studies. A quota-based subsampling strategy, based on journal publication volume, was used to select approximately 100 studies. All four eligible articles from the Korean Journal of Radiology (KJR) were additionally included as a benchmark. Adherence to the 2025 update of MI-CLEAR-LLM was scored through a two-round, consensus-based process: an initial assessment by one reviewer followed by a critical re-evaluation by secondary reviewers, with consensus adjudication by an additional reviewer when needed. Between-journal differences were analyzed with the Kruskal-Wallis test, followed by Dunn post hoc pairwise comparisons with Holm-adjusted P-values.
Results
Of 201 eligible studies identified, 102 were finally analyzed after applying the subsampling strategy. Overall adherence to MI-CLEAR-LLM was moderate (mean, 51.2% ± 14.7%; range, 22.2%-84.2%). Adherence was highest for input data type (100%), test-data independence (80.2%), and adaptation strategy (78.1%), and lowest for prompt execution setup (29.4%) and stochasticity management (33.1%). The least frequently reported items were training-data cutoff date (9.8%) and rationale for prompt wording (15.6%). Adherence varied significantly across journals (P = 0.011), with KJR showing the highest mean adherence (72.8% ± 2.7%).
Conclusion
Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements. Broader adoption of reporting standards is essential to improve the reproducibility and interpretability of future accuracy evaluations.
It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.
Euijun Yang, S. Ko, Hyekyung Woo· Journal of Medical Internet...· 0 citations
Aims: To systematically review large language model (LLM) and natural language processing (NLP) studies published in first-quartile (Q1) clinical radiology journals, focusing on methodological quality, model implementation, and comparative performance. Methods: A systematic search of PubMed and Scopus was conducted to identify original studies involving LLMs or transformer-based NLP systems published in Q1 clinical radiology journals through June 20, 2025. Eligible studies were screened and assessed for methodological characteristics, including dataset type, involving imaging modality (if any), model used, model accessibility, prompt disclosure, and handling of stochasticity. Human-LLM/NLP and LLM/NLP-LLM/NLP performance comparisons were extracted. Results: Fifty-six studies were included, most published in 2024-2025. Proprietary models such as GPT-4 and GPT-4o were most frequently evaluated. Real-world clinical data were used in 62.5% of studies, but only 10.7% reported a power analysis, and 39.1% addressed stochasticity. Prompt engineering was reported in 41.9% of studies. In 455 human-LLM/NLP comparisons, LLMs/NLPs outperformed humans in 54 cases, while humans outperformed in 79; most results (70.8%) were ties. Among 3,164 valid LLM/NLP-LLM/NLP comparisons, GPT-4o had better performance than earlier models. Conclusion: LLMs/NLPs demonstrated performance comparable to radiologists in many text-based tasks but remain inconsistently evaluated. Methodological limitations, including lack of power analysis, incomplete reporting, and under-addressed stochasticity, hinder robust assessment. Greater transparency, standardized evaluation protocols, and inclusion of diverse clinical settings are essential for reliable integration into radiology practice.
I. Mese· Journal of Health Sciences a...· 0 citations
While unsuitable to be used as a sole assessor, ChatGPT-o3 may serve as an adjunct tool to enhance the efficiency and consistency of RoB assessments in systematic reviews.
Siddharth Gandhi, A. Shokravi, Y. Chelliahpillai et al.· PLoS ONE· 0 citations
Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.
Clémentine Bleuze, Karen Fort, Vincent P. Martin et al.· JMIR AI· 0 citations
Abstract Background Large language models (LLMs) show promise in automatically detecting errors in radiology reports, but their performance remains insufficiently validated in large-scale, real-world clinical datasets. Objective This study aimed to systematically evaluate the performance of LLMs in detecting and correcting errors in Chinese radiology reports derived from authentic clinical data. Methods A large-scale dataset of 4480 Chinese radiology reports with modification records containing real clinical practice-generated errors was retrospectively collected between January 2023 and June 2024 at a single institution. After exclusions, 1363 reports containing 1551 errors were included. The dataset covers various anatomical parts of the body from different imaging modalities and was randomly divided into a test set (n=1263) and an internal validation set (n=100). Additionally, 100 error-free reports were added to the internal validation set. An additional 200 English-language reports from the Medical Information Mart for Intensive Care (MIMIC-III) were used for external validation. Eight human readers and 8 widely adopted LLMs, enhanced by prompt engineering, were tasked with error detection. Overall and subgroup detection performance and reading time were evaluated. Correction suggestions from the 2 best-performing LLMs were reviewed by a senior radiologist. Results On the test set, DeepSeek-R1 achieved the highest overall detection rate at 89% (95% CI 87%-90%), significantly better than the other 7 models (P=.001-.007). On the internal validation set, DeepSeek-R1 and Claude-3.5-Sonnet achieved detection rates of 83% (100/120; 95% CI 76%-89%) and 80% (96/120; 95% CI 72%-86%), respectively. DeepSeek-R1 showed performance comparable to radiologists (83%, 95% CI 76%-89% vs 80%, 95% CI 72%-86% for junior radiologists and 78%, 95% CI 70%-85% for senior radiologists; P=.39 and P=.19, respectively) and significantly better performance than that of nonradiologists and nonphysicians (83%, 95% CI 76%-89% vs 66%, 95% CI 57%-74% and 38%, 95% CI 30%-47%; P<.001, respectively). DeepSeek-R1 showed a false-positive rate comparable to radiologists (DeepSeek-R1 vs senior radiologists and junior radiologists, 3% vs 0% and 1%; P=.25 and P=.61, respectively) and a significantly lower rate than nonradiologists and nonphysicians (3% vs 13% and 17%; P=.02 and P=.002, respectively). On the external validation set, DeepSeek-R1 and Claude-3.5-Sonnet achieved detection rates of 94% (95% CI 89%-97%) and 93% (95% CI 88%-97%), respectively. The correction accuracy of DeepSeek-R1 and Claude-3.5-Sonnet was 95% and 91%, respectively. Conclusions Enhanced LLMs, particularly DeepSeek-R1, demonstrated robust performance in error detection and correction within real-world Chinese radiology reports, supporting their clinical use for automated quality assurance and integration into workflows to improve reporting accuracy and efficiency.
Jiafeng Zhou, Yuxin Wei, Qian Cai et al.· Journal of Medical Internet...· 0 citations
AIM
To identify and examine evaluation methods and metrics used for specialist cancer nursing.
DESIGN
A scoping review of published and grey literature on evaluation approaches for specialist cancer nursing roles and models of care.
METHODS
Comprehensive searches were conducted across CINAHL, Cochrane Library, Medline, PsycINFO and Google Scholar for English-language published and grey literature published between January 2014 and November 2025. Two reviewers independently screened and extracted data. Findings were synthesised narratively and mapped to the Strong Model of Advanced Practice Nursing and Quintuple Aim.
RESULTS
Of 3360 records screened, 23 sources met the inclusion criteria: 14 published articles, and 9 grey literature sources (conference abstracts, theses, textbooks). Most sources originated from the USA (n = 12, 52%) or high-income countries (n = 22, 96%), and focused on nurse navigator roles (n = 9, 39%). Five themes emerged in the sources: (1) purpose of evaluation; (2) development of methods and metrics; (3) selection and implementation; (4) data collection approaches; and (5) challenges and considerations. Evaluation was primarily used to demonstrate value and drive quality improvement through pragmatic methods. Metrics varied widely and were concentrated in the Strong Model domains of Direct Comprehensive Care and Support of Systems, with fewer addressing Education, Research and Professional Leadership. Key challenges to evaluation included role variability and lack of standardised tools.
CONCLUSION/IMPLICATIONS
Despite the lack of standardised evaluation practices for specialist cancer nursing, the five themes synthesised in this review can guide evaluation of specialist cancer nursing roles and models of care in real-world settings. Opportunity exists for international collaboration to develop a comprehensive, context-sensitive set of metrics, relevant in diverse healthcare settings, that capture both excellence in service delivery and nursing scholarship.
IMPACT
What problem did the review address? ○ Specialist cancer nurses perform a diverse range of interventions and roles that are complex in nature, leading to challenges in their accurate evaluation. ○ Effective and efficient approaches to evaluation of specialist cancer nursing roles are crucial to demonstrate their value. ○ A significant body of literature has demonstrated the efficacy of specialist cancer nursing roles and models of care in a research framework; however, evaluation is needed to better understand the impact of translating this evidence into real-world settings. What were the main findings? ○ A scoping review exploring evaluation methods and metrics of specialist cancer nursing revealed five key themes: (1) purpose of evaluation; (2) development of evaluation methods and metrics; (3) selection and implementation of evaluation methods and metrics; (4) methods of data collection; and (5) challenges and considerations. ○ Evaluation of specialist cancer nursing is important, however variation in nurses' roles and responsibilities and lack of standardised measurement tools were key challenges. ○ Evaluation metrics varied widely and were specific to specialist cancer nursing roles; predominantly reported under the domains of Direct Comprehensive Care and Support of Systems, with fewer reported under Education, Research and Professional Leadership. Where and on whom will the research have an impact? ○ Nursing and health service leaders can use the identified themes and subthemes as a framework to guide evaluation of specialist cancer nursing. The predominance of English-language and high-income country evidence limits the global applicability of these findings. ○ Gaps in knowledge can drive the future work of cancer nursing organisations to collaboratively develop a comprehensive list of evaluation metrics that can be contextualised for specific roles across diverse health care settings. ○ Specialist cancer nurses in all roles and models of care should have metrics for Research, Leadership and Professional Leadership to support the scholarship of nursing. ○ Consistent national role definitions, shared competency frameworks and standardised outcome measures should be embedded in policy and commissioning to enable systematic evaluation, appropriate resourcing, and integration into workforce and service planning of specialist cancer nurses.
REPORTING METHOD
Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews PRISMA-ScR checklist.
PATIENT OR PUBLIC INVOLVEMENT
Employees of a cancer patient advocacy group were involved in the design of the study, interpretation of the data and the preparation of the manuscript. No patients were involved in the conduct of this scoping review.
Elise Button, Carla Thamm, Megan Crichton et al.· Journal of Advanced Nursing· 0 citations