Skip to content
Review Open access

Large language models for breast cancer treatment planning: a blinded real-world evaluation of DeepSeek, ChatGPT, and oncologist recommendations

Jun 2026 · Frontiers in Digital Health · Vol 8 · 0 citations · 40 references
Medicine

TL;DR

Advanced LLMs, particularly DeepSeek V3.1, demonstrated strong performance in generating standardized, guideline-concordant breast cancer treatment plans, showing superior consistency over human specialists in protocol-driven scenarios, but the widening gap in complex late-stage cases highlights limitations in accounting for clinical context and socioeconomic factors.

Abstract

Rationale and objectives Large language model (LLM) are increasingly explored for oncology decision support, yet their alignment with real-world clinical practice across varying disease complexities remains insufficiently characterized. This study aimed to evaluate and compare the accuracy, stability, and concordance of two advanced LLMs—DeepSeek V3.1 and ChatGPT-5—against experienced oncologists in generating breast cancer treatment plans within a specific clinical setting. Materials and methods This retrospective study compared the performance of DeepSeek V3.1 and ChatGPT-5 with senior oncologists using de-identified records from 213 breast cancer patients (Stages I–IV). To assess effectiveness, we implemented a multidimensional evaluation framework: accuracy was measured using a 5-point Likert scale by three independent, blinded expert reviewers; internal consistency was quantified via variance and coefficient of variation; and clinical concordance was evaluated using a structured five-level scoring system. Statistical analyses, including ANOVA and ordinal regression, were used to examine the impact of disease stage on AI-human agreement. Results Under standardized retrospective review conditions, LLM-generated recommendations demonstrated higher expert-rated guideline concordance and lower variability than historical real-world oncologist plans. Specifically, DeepSeek V3.1 achieved the highest expert-rated accuracy scores with minimal internal variance (4.91 ± 0.36), outperforming both ChatGPT-5 (4.65 ± 0.62) and clinicians (3.82 ± 0.63, P < 0.001). While AI outputs exhibited high mutual consistency (74.2%), expert evaluations revealed a significant decline in AI-clinician agreement as disease stage advanced (P < 0.001), particularly in Stage IV cases where clinicians prioritized real-world constraints such as financial toxicity. Conclusions Advanced LLMs, particularly DeepSeek V3.1, demonstrated strong performance in generating standardized, guideline-concordant breast cancer treatment plans, showing superior consistency over human specialists in protocol-driven scenarios. However, the widening gap in complex late-stage cases highlights limitations in accounting for clinical context and socioeconomic factors. These findings support the role of LLMs as robust clinician-supervised decision-support tools while emphasizing the necessity of human judgment for individualized care.

Read PDF

Similar papers

Jul 2026

Strategies for Deploying Large Language Models for Ascertaining Clinical Outcomes and Sites of Metastases From Radiology Impressions in Patients With Cancer.

PURPOSE To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes. METHODS Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar's odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139. RESULTS This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions). CONCLUSION Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.

Syed Arsalan Ahmed Naqvi, I. Riaz, Amir Saeidi et al. · 0 citations
Review Aug 2026

Large language model treatment-pathway outputs based on structured clinical text in mid- and low rectal cancer: Concordance with multidisciplinary team decisions and features associated with discordance.

Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs, and GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited.

Yuanze Wei, Yulong Tian, Xiaodong Liu et al. · 0 citations
Open access Aug 2026

Simultaneous Preoperative Prediction of Locally Advanced Breast Cancer, DCIS Component, and Multifocality Using Structured Mammographic Features and Gradient-Boosting Machine Learning

Background/Objectives: Accurate preoperative detection of locally advanced breast cancer is essential for neoadjuvant therapy planning. We developed and validated gradient-boosting models using structured BI-RADS mammographic features to simultaneously predict locally advanced breast cancer (LABC), DCIS component, and multifocality in a multi-center cohort. Methods: This retrospective study enrolled 2295 patients from three university-affiliated hospitals; features were coded according to BI-RADS. CatBoost and logistic regression models were built using stratified 60/20/20 splits, with performance assessed via bootstrap resampling, nested cross-validation, and sensitivity analyses. AUROC, AUPRC, Brier score, and calibration metrics assessed discrimination and clinical utility; a leakage audit and SHAP analysis supported interpretation. Results: CatBoost achieved an AUROC of 0.906 (95% CI: 0.876–0.932) for LABC. Because several top predictors overlap with the anatomical criteria defining this outcome, we repeated the analysis excluding them; the reduced model retained a mean AUROC of 0.739, indicating genuine predictive signal beyond the staging overlap. Net benefit was positive across all relevant thresholds, with calibration error of 0.053. DCIS prediction was highly accurate (AUROC 0.979; nested AUROC 0.9707), with no evidence of leakage. Multifocality prediction was more modest (AUROC 0.810), reflecting known limits of two-dimensional mammography. Sensitivity analyses confirmed stable performance across splits, training sizes, and class-weighting schemes. Conclusions: Structured mammographic features combined with gradient-boosting support clinically meaningful, though partly overlapping, risk stratification for LABC; once accounted for, the model still retains independent value. The DCIS model performed very well; multifocality prediction remains more limited, and external validation is needed before clinical use.

Unknown authors · 0 citations
Open access Aug 2026

Augmenting Head and Neck Multidisciplinary Tumor Board Recommendations With Locally Run Large Language Models: Prospective Evaluation of Real-World Implementation

Abstract Background Multidisciplinary tumor boards (MDTs) constitute the foundation of modern tumor therapy. Large language models (LLMs) are widely discussed for optimizing their recommendations. Objective This is the first prospective feasibility study evaluating the implementation of locally run LLMs on real-world cases within a regular head and neck MDT. Methods Seventeen patients participated in the study. The MDT cases were processed by 2 different local LLMs (gemma-3-12b and gpt-oss-20b) to obtain treatment recommendations. The MDT conferred as usual. After the decision was made, the MDT was presented with the LLMs’ recommendations. If deemed to be beneficial, the MDT’s recommendation was adjusted. The MDT members rated the LLMs’ responses inter alia, for medical adequacy on a 6-point Likert scale. In addition, a tabular comparison of the MDT’s and LLMs’ recommendations was carried out. Results In one case, 6% (1/17, 95% CI 0%‐29%), the LLM was able to substantially improve the MDT recommendation by underscoring a follow-up examination that had not yet been performed. Concordance regarding the curative or palliative therapy regimen reached 94% (16/17, 95% CI 71%‐100%); for gemma-3-12b and 59% (10/17, 95% CI 33%‐82%) for gpt-oss-20b. Gemma-3-12b stated the same first-line therapy regimen as the MDT as first-line in 35% (6/17, 95% CI 14%‐62%) of cases, and gpt-oss-20b in 41% (7/17, 95% CI 18%‐67%) of cases. In 59% (10/17, 95% CI 33%‐82%) of patients, gemma-3-12b stated the MDT’s first-line therapy regimen, albeit with a different priority, while for gpt-oss-20b, it was 41% (7/17, 95% CI 18%‐67%) of patients. Medical adequacy, as rated by the MDT members, revealed a median of 5 (IQR 2‐5) for gemma-3-12b and 4 (IQR 3‐5) for gpt-oss-20b. MDT members stated potentially hazardous information in 27% (25/93, 95% CI 18%‐37%) of ratings for gemma-3-12b and 17% (14/83, 95% CI 9%‐26%) of ratings for gpt-oss-20b. Conclusions Locally run LLMs improved the MDT recommendation in 1 case and were primarily useful for identifying potentially relevant missing information in other cases, underscoring that they cannot replace MDTs. However, their observed benefit suggests that more advanced local models may offer safe, rapid, and cost-effective support for MDT decision-making. The study should be seen as an exploratory setting focusing on practical insights rather than benchmarking its clinical impact. Accordingly, the data demonstrate that the integration of LLMs in today’s MDT workflow is feasible and may benefit the quality of decision-making in specific cases.

C. Buhr, L. Müller, Daniel Pinto dos Santos et al. · 0 citations
Open access Jul 2026

A multidimensional benchmarking framework for large language models in oncologic decision making

A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.

M. Halıcı, Serkan Saltürk, Irem Sayin et al. · 0 citations