Aug 2026· Neuro-Oncology Advances· Vol 8· 0 citations
TL;DR
Although contemporary LLMs increasingly reflect medical consensus for CNS metastases, inconsistent reliability remains a concern, underscoring the need for caution in patient use.
Abstract
Abstract Public-facing large language models (LLMs) are increasingly used by patients to obtain medical information, yet their accuracy, safety, and readability in the context of central nervous system (CNS) metastases remain poorly characterized. Fifteen simulated patient questions regarding CNS brain metastases were submitted to four LLMs (ChatGPT, Claude, Gemini, and Open Evidence). Responses were evaluated using a 5-point Likert scale for accuracy based on National Comprehensive Cancer Network (NCCN) guidelines. Incorrect responses were defined as scores ≤2. Readability was assessed using Flesch Reading Ease (FRE), Flesch–Kincaid Grade Level (FKGL), and Gunning Fog index. Word count was also analyzed. Mixed-effects models were used to compare performance across models, accounting for repeated measures by question. ChatGPT demonstrated the highest mean accuracy score (4.8), followed by Open Evidence (4.7), Gemini (4.1), and Claude (3.7). Incorrect responses occurred most frequently with Claude (26.7%), followed by Gemini (13.3%), while no incorrect responses were observed for ChatGPT or Open Evidence. In ordinal regression analysis, ChatGPT and Open Evidence demonstrated superior performance compared to Claude and Gemini (p < 0.01), with no significant difference between ChatGPT and Open Evidence. Readability analysis revealed that Gemini produced the most readable responses (FRE 43.8; FKGL 12.3; Gunning Fog 15.0), followed by ChatGPT, while Claude and Open Evidence generated significantly less readable outputs (p < 0.01). Gemini also generated the longest responses (mean 467 words), whereas ChatGPT and Open Evidence produced shorter responses (∼370 words). LLM performance varied substantially across accuracy, safety, and readability. ChatGPT and Open Evidence achieved the highest accuracy with no incorrect responses, whereas other models, despite greater readability, were more likely to generate incorrect information. Although contemporary LLMs increasingly reflect medical consensus for CNS metastases, inconsistent reliability remains a concern, underscoring the need for caution in patient use.
PURPOSE
To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes.
METHODS
Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar's odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139.
RESULTS
This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions).
CONCLUSION
Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.
Syed Arsalan Ahmed Naqvi, I. Riaz, Amir Saeidi et al.· JCO Clinical Cancer Informat...· 0 citations
Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.
Jiwon You, Hangsik Shin· npj Digital Medicine· 0 citations
To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible large language model chatbots when answering patient-facing lung cancer prognostic questions under standardized single-turn English prompting.
In this Chatbot Health Advice Reporting Transparency-guided cross-sectional evaluation, 53 standardized English prompts were submitted once to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao through official web interfaces during April 1–21, 2026. Five blinded raters assessed 265 responses for safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, Global Quality Scale, and readability. Paired repeated-measures analyses were used.
Inter-rater agreement was good to excellent. Safety differed significantly across models (Cochran’s Q = 14.089, df = 4,
p
= 0.007). Gemini generated the highest proportion of safe responses (48/53, 90.6%), whereas DeepSeek generated the lowest (33/53, 62.3%). The only adjusted pairwise safety difference that remained significant was Gemini versus DeepSeek (adjusted
p
= 0.023). Accuracy, empathy, reliability, information quality, and readability also differed significantly across models (all
p
< 0.001). Gemini showed the most favorable descriptive profile for safety, accuracy, empathy, and reliability, while Copilot produced the most readable responses.
Public-facing chatbots differed substantially in safety, reliability, communication quality, and readability. These findings are time-, interface-, and prompt-dependent. Chatbots may support general patient education but should not replace individualized clinician-led prognostic communication.
Yan-Ru Jiang, Qianyun Wang, Liang Zheng et al.· Frontiers in Public Health· 0 citations
A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.
M. Halıcı, Serkan Saltürk, Irem Sayin et al.· Scientific Reports· 0 citations
Background As society is increasingly depending on large language models (LLMs) for health-related questions, it is essential to objectively evaluate the quality and accessibility of the oral cancer information they provide. Although LLMs occupy a growing space in digital health communication, it remains unknown whether the information they generate is both reliable and easy to read. Objective The purpose of this study was to evaluate the reliability and readability, respectively, of the responses generated by four mainstream LLMs (ChatGPT, Gemini, Perplexity, and DeepSeek) to questions related to oral cancer. Specifically, the present authors aimed to evaluate the reliability of responses to common oral cancer questions and assess whether the readability of responses meets established expectations. Methods Twenty-two commonly asked, patient-orientated oral cancer–related questions were developed through two predefined phases: Google Trends analysis and expert consultation with specialists in oral oncology. Each question was entered as an independent single-turn prompt into four LLMs: ChatGPT-5, Gemini 2.5, Perplexity Pro, and DeepSeek v3.2. The primary outcome was information reliability and quality, assessed using four standardized instruments: the DISCERN questionnaire, the Ensuring Quality Information for Patients (EQIP) tool, the Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Scale (GQS). The secondary outcome was readability, assessed using six established indices: the Automated Readability Index, Flesch Reading Ease Score, Gunning Fog Index, Flesch-Kincaid Grade Level, Coleman-Liau Index, and Simple Measure of Gobbledygook. Results Significant differences were observed among the four LLMs in DISCERN, EQIP, and JAMA scores (all P < 0.001), whereas no significant difference was found in GQS scores (P = 0.440). Perplexity Pro achieved the highest mean DISCERN score (46.36 ± 4.70), EQIP score (85.00 ± 0.00), GQS score (4.05 ± 0.58), and JAMA score (1.00 ± 0.00). However, all models produced responses above the recommended sixth-grade readability level. The mean FKGL scores ranged from 12.65 ± 3.07 for ChatGPT-5 to 15.65 ± 3.36 for Perplexity Pro, and the mean FRES scores ranged from 37.50 ± 14.27 for Perplexity Pro to 52.59 ± 12.47 for Gemini 2.5. Conclusion Current LLMs may support oral cancer patient education, but their use remains limited by variable information quality, insufficient transparency, and poor readability. Although Perplexity Pro performed better on several reliability-related metrics, no model showed consistently high performance across all dimensions or met recommended readability standards. Future LLM-based patient education tools should prioritise verifiable sourcing, guideline-based accuracy, risk communication, and plain-language adaptation.
Bo Zhang, Weidi Shi, Ying Zhang· Oral Health & Preventive Den...· 0 citations
This study investigates the readability, clinical reliability, and temporal consistency of artificial intelligence (AI) chatbots regarding pneumothorax information. A question bank comprising 40 patient-centered queries was deployed across three large language models (ChatGPT, Gemini, Copilot), stratified by two access tiers and two prompting strategies (zero-shot versus the optimized PROMPORT strategy). Queries were replicated longitudinally on Days 1, 3, and 7 under strict session-control protocols. Text accessibility was quantified using five automated readability indices, while two independent, blinded thoracic surgeons evaluated clinical quality using modified DISCERN (mDISCERN), JAMA benchmarks, and PEMAT-P indices. Readability metrics demonstrated absolute structural stability across the tracking intervals (p > 0.05). Unprompted configurations consistently generated complex, high-school-level outputs, whereas the PROMPORT strategy successfully compressed linguistic variances and neutralized chronological algorithmic drift (p > 0.05). Conversely, unprompted architectures exhibited significant temporal volatility in mDISCERN and JAMA profiles (p < 0.05), which was successfully stabilized by optimized prompt constraints. Inter-rater reliability was high across all structural evaluations. In conclusion, while unprompted models exhibit marked baseline linguistic and quality variations, the strategic integration of robust prompt engineering successfully enforces the temporal stability and clarity required for reliable digital public health communication.
Ömer Önal, Suzan Temiz Bekce· Scientific Reports· 0 citations