Skip to content

Patient preference for large language model (LLM)–optimised oncology trial information: A randomised controlled crossover study.

Jul 2026 · Journal of Clinical Oncology · Vol 44, pp. 6-6 · 0 citations

TL;DR

LLM optimisation improves patient preference for trial information compared with standard registry descriptors, supporting further evaluation of its use in rendering patient-facing materials.

Abstract

6 Background: Clinical trial participation in oncology is frequently hindered by complex, jargon-laden information that impairs comprehension by patients and caregivers. LLMs offer a scalable solution to simplify technical text, but their utility in oncology has not been evaluated. Methods: We conducted a randomised, controlled, three-period crossover study at two tertiary breast cancer outpatient clinics in Sydney, Australia. Eligible patients and caregivers evaluated trial descriptions across 3 formats: standard ClinicalTrials.gov text (Control); LLM-optimised text generated using zero-shot prompting with GPT-4 (LLM); and further refined by an oncologist (LLM+E). Latin square randomisation controlled for order effects. The primary endpoint was the Global Preference Score (GPS), operationalised as the minimum Likert rating (1-5 scale) across 5 sections (Title, Summary, Intervention, Description, Eligibility) to reflect that incomprehensibility of any component degrades the overall document utility. Sample size (n≥18) was determined via Monte Carlo simulation to detect a minimum 1 point Likert shift with 90% power (α=0.05); the recruitment target was 36 (accounting for 50% attrition). Primary analysis employed Friedman rank sum test blocked by participant, with a cumulative link mixed model (CLMM) including random intercepts for participants to adjust for age, education, and first language. All LLM generated text (LLM±E) were vetted for accuracy of content by ≥1 medical oncologist. Results: Between September and December 2025, 30 of 31 recruited participants provided crossover responses for primary analysis (401 valid responses across 5 sections, 11% invalid). The mean age was 53 years (95% CI, 44-63), with the majority of respondents being female (n = 27, 90%), native English speakers (n = 24, 80%) and holders of tertiary qualifications (n = 17, 57%); 20 respondents were patients (67%). The median GPS was 3.0 (IQR 2.5) for Control, 4.0 (IQR 2.0) for LLM and 3.0 (IQR 1.5) for LLM+E. LLM achieved a significantly higher GPS vs. Control (p = 0.02, Friedman test). Multivariable CLMM confirmed increased odds of higher preference ratings for LLM text vs. Control (OR 1.83, 95% CI 1.01–3.35, adjusted p = 0.048). Paradoxically, LLM+E showed no improvement over Control (padj = 0.826, r = 0.063). Native English speakers rated contents more critically than non-native speakers (OR 0.092, p = 0.009), independent of arm allocation. Conclusions: LLM optimisation improves patient preference for trial information compared with standard registry descriptors, supporting further evaluation of its use in rendering patient-facing materials. Expert oncologist revision may inadvertently re-introduce complexity, negating the linguistic accessibility gains provided by LLMs, suggesting that human-in-the-loop workflows need caution to preserve linguistic style for accessibility.

View source

Similar papers

Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

Objective: Large language models (LLMs) are increasingly being used to develop medical education materials; however, it remains unclear how reliable, readable, or guideline-compliant the content generated by these models is for non-English-speaking patient groups. We evaluated the quality of Turkish melanoma patient education texts generated by seven frontier LLMs. Methods: A standardized 22-item Turkish prompt, built from international melanoma guidelines, was put to seven models in zero-shot sessions: ChatGPT 4.0 Turbo, Gemini 2.0 Flash, Claude 3.7 Sonnet, Grok 3, Qwen 2.5 Plus, DeepSeek R1, and Mistral Large 2. Each output was rated for readability (Ateşman Index), understandability, how clearly medical terminology was explained, scientific reliability (DISCERN instrument), empathy, and adherence to a 31-item guideline-based checklist. Model comparisons were summarized descriptively, using model-level absolute scores, score ranges, and rankings. Results: Model performance differed across readability, understandability, reliability, empathy, and guideline-adherence domains. DeepSeek R1 led on both readability (81.6) and understandability (23.5/25). Guideline adherence was strongest for Grok 3 and DeepSeek R1, at 96.8% and 93.5%, respectively, and Grok 3, DeepSeek R1, and Gemini 2.0 Flash each scored above 90% on the normalized total DISCERN measure. DeepSeek R1 also recorded the highest empathy score (90%). Gemini 2.0 Flash had the lowest readability score and produced the longest output (Ateşman 65.8). None of the models provided citations or verifiable sources, so every model received the lowest possible DISCERN Source Reliability score; Mistral Large 2 showed the weakest overall performance. Conclusion: How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next. A few produced text that was clear, empathetic, and reasonably guideline-concordant, but the lack of verifiable citations and uneven guideline coverage remain genuine limitations. These findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Niyazi Çetin, A. Atılan · 0 citations
Review Open access Aug 2026

The daily dose: Early usability of an LLM tool for patient summaries and trial matching in radiation oncology

Background Radiation oncology workflows generate large volumes of electronic health record (EHR) data requiring daily synthesis. Large language model (LLM)-based automation is promising, but workflow-embedded implementations at scale remain limited. We describe the design and early usability and adoption evaluation of The Daily Dose (TDD), an LLM-driven system for automated clinical summarization and trial identification in radiation oncology. Materials and methods TDD delivers physician-specific email summaries each morning across three Mayo Clinic campuses using RadOnc-GPT (GPT-4o) to generate EHR-derived patient summaries and identify potentially eligible clinical trials for new or consult visits. One month post-deployment, an anonymous cross-sectional survey adapted from the System Usability Scale and Technology Acceptance Model was administered to all recipients. Results Fifty-five of 110 users responded (50%); 94.5% were in radiation oncology and 69.1% were attending physicians. Overall, 83.6% used TDD at least several times per week. Mean domain scores (5-point Likert) were 3.89 ± 1.04 for usability and satisfaction, 3.43 ± 1.24 for perceived usefulness, and 3.80 ± 1.17 for impact and future use. Satisfaction was significantly associated with perceived time savings (p < 0.001); 27% estimated saving ≥10 min daily. Internal consistency was high (α = 0.97). Free-text responses highlighted improved preparedness and patient-context awareness but noted occasional inaccuracies and imperfect trial matching. Conclusion In this early usability and adoption evaluation, a workflow-integrated LLM summarization tool was widely adopted and generally favorably perceived. These findings reflect user perceptions; objective validation of summary accuracy, trial-matching performance, and workflow efficiency is needed to establish clinical impact.

J. Holmes, F. Mastroleo, M. Borras-Osorio et al. · 0 citations
Open access Jul 2026

Performance of leading large language models in adhering to clinical guidelines for anaplastic thyroid cancer: a comparative study

Leading LLMs show variable capacity to align with ATC clinical guidelines, while top-performing models hold promise as supportive tools, their inconsistencies across domains and complexity levels preclude autonomous clinical use.

Mohamed Yasser, Ghada Barakat, S. Awny et al. · 0 citations
Jul 2026

Evaluating the Accuracy of ChatGPT-4o in Addressing Complex Clinical Questions Based on NCCN Guidelines for Rectal Adenocarcinoma.

ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines across all evaluated clinical domains with notable improvement over prior ChatGPT iterations evaluated by this group.

Ryan J Meyer, Tamir E. Bresler, Kevin Palmer et al. · 0 citations
Open access Jul 2026

Evaluation and comparison of large language model responses to patient questions after diagnosis of high-risk human papillomavirus infection: an expert-rated digital patient education study

Background After diagnosis of high-risk human papillomavirus (HPV) infection, patients often seek guidance on cancer risk, colposcopy, follow-up, treatment, partner management, pregnancy, and anxiety. Large language models (LLMs) are increasingly used for medical consultation, but their response quality and safety in this setting require evaluation. Methods In this single-center, expert-rated methodological study, 15 patient-style questions based on common outpatient consultations were submitted to ChatGPT-5.5 Instant and DeepSeek-V3. Each model generated three independent responses per question, yielding 90 artificial intelligence (AI)-generated responses. Five blinded gynecology experts independently evaluated all responses, producing 450 expert-rating records. Accuracy, safety, guideline concordance, completeness, and understandability were rated on a 1–5 Likert scale. Experts also assessed major error and potential harm. The primary outcome was the proportion of responses without major error or potential harm. Results ChatGPT-5.5 Instant had higher response-level scores than DeepSeek-V3 for accuracy (mean difference, 0.25; 95% CI, 0.11–0.39; FDR-adjusted p = 0.010), safety (0.27; 0.07–0.47; p = 0.010), completeness (0.22; 0.07–0.37; p = 0.020), and composite score (0.18; 0.04–0.31; p = 0.027). The difference in guideline concordance did not remain significant after multiplicity correction (0.16; −0.01 to 0.33; FDR-adjusted p = 0.057), and understandability was similar. Responses without major error or potential harm occurred in 42/45 (93.3%; 95% CI, 81.7–98.6%) ChatGPT responses and 38/45 (84.4%; 70.5–93.5%) DeepSeek-V3 responses (risk difference, 8.9 percentage points; 95% CI, −13.3 to 31.1; p = 0.344). Exploratory question-level risks were observed in HPV16/18 positivity, normal cytology, partner management, and pregnancy scenarios. Conclusion Both two models demonstrated generally strong performance across the evaluated quality domains. ChatGPT achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant. Both models require guideline-based clinical oversight.

Zhen Hao, Lin Wang, Yue Wu et al. · 0 citations
Review Aug 2026

Large language model treatment-pathway outputs based on structured clinical text in mid- and low rectal cancer: Concordance with multidisciplinary team decisions and features associated with discordance.

Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs, and GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited.

Yuanze Wei, Yulong Tian, Xiaodong Liu et al. · 0 citations