Skip to content
Review Open access

The Role of Advanced Language Models in MedicalDiagnostics: A Case Study on Breast Cancer Prediction

Aug 2026 · International Journal of Interactive Multimedia and Artificial Intelligence · 0 citations

TL;DR

Although the evaluated LLMs did not outperform traditional supervised models, the study provides a clear performance baseline for future research on structured clinical prediction with language models and shows that LLMs may offer value as complementary exploratory tools, but their outputs should be interpreted only with expert oversight.

Abstract

Breast cancer remains one of the most common and life threatening cancers worldwide, and early detection is strongly associated with improved survival and reduced treatment burden. This study  investigates the ability of Large Language Models to perform diagnostic prediction from structured breast cancer related data. We systematically evaluated 12 LLMs across three public datasets with different clinical characteristics: the Wisconsin Breast Cancer Dataset (WBCD) based on cytological features, the Breast Cancer Coimbra Dataset (BCCD) based on metabolic biomarkers, and the Mammographic Mass Dataset (MMD) based on mammographic attributes. The evaluation covered 13 prompting strategies, including three zero-shot variants, few-shot prompting, and multiple Chain-of-Thought (CoT) and knowledge-enhanced reasoning settings. Performance was assessed using confusion-matrix-based metrics, including accuracy, precision, recall, F1-score, specificity, and Matthews Correlation Coefficient. The results showed that performance was strongly dependent on both dataset type and prompting design, and no single model dominated all tasks. The best model–strategy pairvaried by dataset: Cogito-v1-preview-qwen-32B achieved the highest F1-score on WBCD with an F1-score of 92.00%, GPT 4.1 and GPT 4o on BCCD with an F1-score of 85.39%, and Gemini 2.5 Flash Lite on MMD with an F1-score of 82.91%. Prompt engineering had a substantial effect on outcomes, but its benefit varied across models, with some systems improving under knowledge-enhanced few shot prompting and others performing best under simpler strategies. Comparison with state-of-the-art traditional ML baselines showed that, while LLMs do not yet surpass supervised methods, the performance gap has narrowed substantially, particularly on MMD, where the best single-run gap was 2.04 percentage points in F1 and the mean gap under robustness analysis was approximately 5.3 points. Robustness analysis across multiple few-shot example sets confirmed stable performance on WBCD (F1 = 91.37 ± 0.71%) and BCCD (F1 = 87.46 ± 1.91%), while revealing moderate sensitivity on MMD (F1 = 79.62 ± 2.87%). Although the evaluated LLMs did not outperform traditional supervised models, the study provides a clear performance baseline for future research on structured clinical prediction with language models. The results show that LLMs may offer value as complementary exploratory tools, but their outputs should be interpreted only with expert oversight because clinically significant errors remain.

Read PDF

Similar papers

Open access Aug 2026

Breast Cancer Risk Prediction and IDC Tumor Classification in Different Stages Based on Machine Learning Models

A multi-modal, data-driven intelligent framework especially designed to breast cancer analysis the Multi-Dimensional Feature Refinement Cancer Network (MDFR-Cancer-Net) that can be used to analyse breast cancer at various clinical stages based on multi-modal data is proposed.

Yun-Ze Li, N. Sani · 0 citations
Open access Aug 2026

An Interpretable Machine Learning Framework for Breast Cancer Diagnosis Using Statistical Feature Analysis and Ensemble Classification

A clear, statistically sound, yet easily understandable breast cancer diagnosis is a difficult issue in all healthcare systems, because early stages of breast cancer are critical in therapy success and long-term survivability. This machine-learning-based breast cancer classifier, in a statistically justified, rigorousl...

T. Haripriya, M. V. Ramana Murthy, Ch. Vasavi et al. · 0 citations
Open access Sep 2026

Machine Learning and Explainable AI for Breast Cancer Patient Prioritization: An Intelligent Decision-Support Framework

The proposed framework can serve as an intelligent decision-support tool for prioritizing breast cancer patients and improving resource allocation when healthcare capacity is constrained and its relatively simple and scalable architecture facilitates potential implementation in healthcare environments with limited reso...

Fabián Silva-Aravena, J. Morales, Hugo Núñez Delafuente et al. · 0 citations
Open access Sep 2026

High-Dimensional Breast Cancer Classification: Evaluating Trade-offs Between Accuracy, Robustness, and Computational Cost

Background: Breast cancer is the leading cause of cancer-related mortality in females, with 2.3 million new cases diagnosed annually. Machine learning (ML) algorithms have the potential to improve diagnostic accuracy through the analysis of high-dimensional data. However, the lack of standardized benchmarks across diff...

Ali A. Hamad, Naaman Omar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.