Leading LLMs show variable capacity to align with ATC clinical guidelines, while top-performing models hold promise as supportive tools, their inconsistencies across domains and complexity levels preclude autonomous clinical use.
Abstract
Anaplastic thyroid cancer (ATC) is a rare, aggressive malignancy with poor prognosis. Adherence to guidelines from the National Comprehensive Cancer Network (NCCN), American Thyroid Association (ATA), and European Society for Medical Oncology (ESMO) is critical for optimal patient outcomes. As large language models (LLMs) increasingly enter clinical workflows, rigorous evaluation of their alignment with established guidelines is essential. We evaluated five leading LLMs for their ability to generate guideline-concordant responses to clinical questions about ATC. We conducted a comparative study in 2025 following TRIPOD-LLM (Transparent Reporting of a Multivariable Model for Individual Prognosis or Diagnosis, Large Language Models) guidelines. Seventy clinical questions of varying complexity were developed from ATA, NCCN, and ESMO guidelines. Three surgical oncology experts validated each question and subsequently evaluated responses from five LLMs: ChatGPT 4.1, ChatGPT 5, Gemini 2.5 Pro, Claude Sonnet 4, and DeepSeek R1. Each response was scored for relevancy, clarity, accuracy, and adequacy on a 5-point Likert scale. Inter-rater reliability was assessed using both intraclass correlation coefficients (ICC) and Gwet’s AC2 with ordinal weights. Model comparisons used the Kruskal-Wallis test with Dunn’s post-hoc analysis and Bonferroni correction. A pre-specified sensitivity analysis excluding the unblinded model (ChatGPT 5) was performed to confirm robustness. Significant performance differences emerged across all four metrics: accuracy (p = 0.007), adequacy (p = 0.003), clarity (p = 0.014), and relevance (p < 0.001). Gemini 2.5 Pro achieved the highest median accuracy (4.5), followed by DeepSeek R1 (4.4), while ChatGPT 4.1 scored lowest (4.0). ICC values ranged from 0.34 to 0.44 (poor to moderate), but Gwet’s AC2 yielded substantially higher estimates of 0.61 to 0.73 (moderate to substantial agreement), reflecting the impact of restricted score range on conventional reliability metrics. The sensitivity analysis excluding ChatGPT 5 confirmed the performance hierarchy among blinded models, with significance preserved or strengthened across all four metrics. Leading LLMs show variable capacity to align with ATC clinical guidelines. While top-performing models hold promise as supportive tools, their inconsistencies across domains and complexity levels preclude autonomous clinical use. These models should serve strictly as decision aids under expert supervision.
Abstract Background Ovarian cancer (OC) is a highly fatal gynecologic malignancy with complex management challenges and limited long-term survival for advanced stages. Large language models (LLMs)—including systems such as GPT-4, Claude, Google Gemini, and others—are emerging artificial intelligence (AI) tools capable of performing health care–related tasks such as diagnostic support, treatment planning, report generation, and patient communication. However, their applications in OC care have not yet been comprehensively assessed. Objective This protocol outlines a systematic review and meta-analysis aimed at evaluating the use, performance, and clinical impact of LLMs in OC management. We will examine how LLMs have been applied across various domains (eg, diagnosis, prognosis, treatment planning, and patient engagement), the metrics used to assess their performance (eg, accuracy, sensitivity, and area under the curve), and their strengths and limitations. Methods This review will be conducted in accordance with PRISMA-P (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols) guidelines. A comprehensive search strategy will be implemented across biomedical, technical, and Chinese-language databases (eg, PubMed, Embase, Web of Science, IEEE Xplore, and China National Knowledge Infrastructure) from inception to December 31, 2025. Eligible studies include clinical evaluations, validation studies, and real-world implementation reports involving LLMs in OC care. Two independent reviewers will perform screening, data extraction, and quality appraisal using validated tools (eg, version 2 of the Cochrane risk-of-bias tool for randomized trials, Risk of Bias in Nonrandomized Studies of Interventions, Quality Assessment of Diagnostic Accuracy Studies 2, and Prediction Model Study Risk of Bias Assessment Tool+AI). Outcomes of interest include model performance metrics, clinical process impacts, safety concerns, and usability. Meta-analyses will be conducted where feasible using random-effects models in R (meta, metafor, and mada packages), including bivariate models for sensitivity and specificity. Results The review is currently in progress. The PROSPERO registration has been completed, and the literature search and selection process is underway. Study selection, data extraction, and quality assessment are expected to be completed by mid-2026. Final results will include pooled performance metrics (eg, accuracy, F1-score, and area under the curve), qualitative insights into clinical integration, and identification of limitations such as reporting bias or insufficient external validation. Conclusions This systematic review will provide the first comprehensive synthesis of evidence on the application of LLMs in OC care. It will identify promising use cases, highlight safety and reporting challenges, and inform future research directions. The findings are expected to support evidence-based integration of LLMs into gynecologic oncology workflows while promoting transparency and methodological rigor in AI evaluation.
Yanhong Wang, Jialiang Yao, Jianhui Tian et al.· JMIR Research Protocols· 0 citations
Abstract Background Maintenance of oncology clinical practice guidelines (CPGs) is increasingly challenged by the rapid growth of trial data and therapeutic complexity. While large language models (LLMs) have shown promise in information retrieval, their utility in the rigorous, end-to-end workflow of guideline maintenance remains underexplored. Objective This case study aimed to systematically evaluate the performance of frontier LLMs in supporting oncology guideline maintenance. We sought to determine their reliability in predicting necessary guideline updates based on new evidence, their accuracy in extracting data from clinical trials, and their effectiveness as automated auditors for detecting errors in established guidelines. Methods Using the Onkopedia peripheral T-cell lymphoma (PTCL) guideline as a prospective case study, we tasked frontier models with deep-research modes and autonomous web-search capabilities (Gemini 2.5 Pro and GPT o4-mini-high) to predict a guideline update in August 2025 based on the 2021 version. Predictions were validated against the official 2025 revision published in October 2025. Next, we benchmarked evidence extraction accuracy across 80 pivotal trials using models of varying scale (27B-671B parameters vs frontier). Finally, we deployed a stacked LLM workflow to audit 28 recently updated Onkopedia guidelines for linguistic and content-related errors. Results In the predictive task, models captured 36.7% to 40% of substantive updates, often identifying landmark approvals, but frequently overstating evidence. An independent, model-blinded rescoring yielded substantial agreement (weighted Cohen κ=0.75) and confirmed predictive accuracies of 35% to 38.3%. While frontier models demonstrated high accuracy (up to 99.2%) in extracting data from individual studies, substantially outperforming smaller open-source models, this precision declined during multisource synthesis. We observed a position-dependent performance drop in long-form generation, with GPT’s endpoint accuracy dropping from 84.6% in the first half of the drafted guideline to 38.5% in the second half. As automated auditors of existing CPGs, the models successfully identified a median of 16.5 (IQR 13.8-20.3) formal errors per document and detected several clinically relevant inconsistencies (eg, invalid scoring formulas and incorrect staging definitions). Conclusions LLMs currently lack the reasoning stability for autonomous guideline authoring due to deficits in complex synthesis. However, they are effective tools for high-fidelity evidence extraction and automated quality assurance, supporting a human-led, AI-augmented workflow for efficient guideline maintenance.
M. Knauer, Julian Greß, J. Kather et al.· JMIR AI· 0 citations
A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.
M. Halıcı, Serkan Saltürk, Irem Sayin et al.· Scientific Reports· 0 citations
Abstract Background Although large language models (LLMs) have demonstrated the ability to generate the impression section from radiology findings automatically, the incremental diagnostic value of clinical information for these models remains unclear. Objective This study aimed to evaluate the incremental diagnostic value of clinical information for LLMs and compare their performance with that of radiologists. Methods This retrospective study included radiology reports from patients with histopathologically confirmed liver, lung, and breast diseases from 2 institutions between October 2021 and February 2025. We defined three progressive information input scenarios: (1) basic patient information and imaging findings, (2) scenario A plus chief complaint or clinical history, and (3) scenario B plus key laboratory results. Scenario-based data were input into 3 general-purpose LLMs (DeepSeek-R1, Gemini 2.5 Pro, and GPT-4o), generating 2709 entries. Diagnostic accuracy was assessed for both benign-malignant differentiation and disease diagnosis, with histopathology serving as the reference standard. Accuracy was compared among scenarios and against radiologist performance using the McNemar test, and P values were adjusted using the Holm-Bonferroni correction for multiple comparisons. Results A total of 301 patients with pathologically confirmed diseases were included (mean age 53.5, SD 12.0 years; women: n=208, 69.1%). In the liver cohort, a numerical trend toward higher accuracy was observed in scenario C compared with scenario A across all 3 models (scenario C range: 72.3%‐76.2% vs scenario A range: 64.4%‐68.3%); these differences did not reach statistical significance after Holm-Bonferroni correction (all adjusted P>.99). Notably, the DeepSeek-R1 model in scenario C achieved the highest diagnostic accuracy (77/101, 76.2%), with no evidence of a difference compared with radiologists (82/101, 81.2%; adjusted P>.99). In contrast, results in the lung and breast cohorts were more heterogeneous. In the lung cohort, GPT-4o achieved its highest accuracy in scenario A for disease diagnosis (68/92, 73.9%), which exceeded its performance in scenario B (64/92, 69.6%) and scenario C (66/92, 71.7%), suggesting that additional clinical information did not confer a consistent benefit. Gemini 2.5 Pro in scenario B achieved the highest accuracy in this cohort (72/92, 78.3%); however, no statistically significant difference was found compared with radiologists (80/92, 87.0%; adjusted P=.25). In the breast cohort, DeepSeek-R1 achieved the numerically highest diagnostic accuracy in scenario A, and it decreased numerically with the addition of laboratory tests, although no significant difference was found between scenarios A and C (73/108, 67.6% vs 71/108, 65.7%; adjusted P>.99). Conclusions While the addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.
Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al.· Journal of Medical Internet...· 0 citations
PURPOSE
To develop a large language model (LLM) (Truveta Language Model Oncology [TLM-Oncology]) to extract real-world oncology staging data across multiple cancer types from clinical documentation with high precision.
METHODS
We selected patients from a large integrated health system with a bladder, cervical, colorectal, breast, or prostate cancer diagnosis in their structured data. We identified relevant notes using note metadata and keywords and annotated overall stage; T, N, and M; associated timeframe; and cancer diagnosis on a sample of 700 notes as ground truth. Of the 700 notes, 450 were divided equally between training, validation, and test sets for bladder, cervical, and colorectal cancers; 150 were used for targeted error-pattern training on these cancers; and the remaining 100 were split equally between breast and prostate cancer test sets. We started with a pretrained LLM and applied supervised fine-tuning to adapt the model to structured clinical information extraction. Model performance was measured using precision, recall, and F1 scores at the relation level and individual attribute level.
RESULTS
We extracted over 2.5 million staging records for 217,768 patients from over two million notes. Relation-level precision across the six attributes ranged from 0.77 to 1.0 for the first three cancers and, without further training, 0.83 to 1.0 for two additional cancers.
CONCLUSION
TLM-Oncology extracted detailed cancer staging information for five cancers from a variety of clinical documentation within a single integrated health system with high precision and turned data that were previously inaccessible into a valuable resource for downstream use. We are currently evaluating TLM-Oncology on other solid tumors within three additional health systems to assess its generalizability.
S. Abhyankar, Rajesh Rao, Mehraveh Salehi et al.· JCO Clinical Cancer Informat...· 0 citations