A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.
Abstract
Large language models (LLMs) are increasingly explored as clinical decision support tools in oncology; however, reliance on isolated metrics has limited the development of multi-dimensional evaluation frameworks. This comparative observational study utilized five stepwise, clinically realistic non-small cell lung cancer scenarios reflecting real-world diagnostic, therapeutic, and follow-up decision-making. Open-ended clinical questions were answered by three LLMs (Gemini 2.5 Pro, GPT-5, and Claude Opus 4.1) via their official APIs and compared with evidence-based reference answers. Model outputs were evaluated using expert-rated clinical accuracy and explainability, alongside operational metrics including cost, response time, and generative efficiency. All dimensions were integrated into an expert-weighted Composite Performance Score (CPS). Across 30 clinical questions, significant inter-model differences were observed for all metrics (p < 0.001). GPT-5 achieved the highest accuracy, explainability, and generative efficiency, while Gemini 2.5 Pro demonstrated the lowest cost and Opus 4.1 the fastest response times. Integrated analysis yielded the highest CPS for GPT-5, followed by Gemini 2.5 Pro and Opus 4.1 (Kendall’s W = 0.87). A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice. Nevertheless, the use of LLMs in this domain should remain clinician-supervised.
Large language models (LLMs) are emerging as tools to support clinical decision making. HIV management is a compelling use case due to its complexity and dynamic nature, involving diverse treatment options, comorbidities, and adherence challenges. However, integrating LLMs into clinical practice raises concerns about accuracy, safety, and clinician acceptance. Despite growing interest, their performance in HIV care remains poorly studied, and benchmarking is lacking.
We developed HIVMedQA, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios. We evaluated seven general-purpose and three medical LLMs. Performance was assessed using lexical similarity and an extended medical LLM-as-a-judge framework capturing key clinical dimensions, including question comprehension, reasoning, knowledge recall, bias, potential harm, and factual accuracy, to better capture nuances relevant to the medical domain, with additional evaluation by HIV-experienced physicians.
Performance varies substantially across models and task complexity. Gemini 2.5 Pro achieves the highest overall scores, followed by Claude 3.5 Sonnet and MedGemma-27B. Knowledge recall is generally stronger than question comprehension or clinical reasoning. Medical LLMs do not consistently outperform general-purpose models, and model size alone does not predict performance. Several models are sensitive to cognitive bias prompts. LLM-as-a-judge scoring aligns better with clinician assessment than lexical metrics.
HIVMedQA provides a structured benchmark for evaluating LLMs in HIV clinical decision support. Current LLMs show promise, but limitations in reasoning, bias robustness, and safety indicate that careful validation, domain-specific evaluation, and clinician oversight remain essential before clinical deployment.
Gonzalo Cardenal-Antolin, J. Fellay, Bashkim Jaha et al.· Communications Medicine· 0 citations
Background The EU4Health project PCM4EU aimed to improve survival rates and quality of life of patients with cancer based on precision cancer medicine. To achieve this, enhanced expertise and quality of molecular cancer diagnostics are key. Clinical decision support systems (CDSS) have become increasingly important after the introduction of comprehensive genomic diagnostic profiling. While external quality assessment schemes are mandatory for most diagnostic laboratory tests, similar programs for CDSS tools are currently lacking. To address this, we piloted an international ring test documenting CDSS usage, performance, and manual interpretation. Materials and methods Twenty synthetic datasets were generated, mimicking small variant call sets from a typical targeted 500-gene panel (VCF format) across multiple cancer types (10 tumour-normal pairs and 10 tumour-only). Participants received standardised instructions via e-mail and at a virtual meeting and submitted results using a structured response form. Results Eight laboratories from seven countries participated. All participants submitted results for the 10 tumour-only cases; one submitted results from two assessors. Tumour-only cases contained 6-18 variants where interpretation could be critical. Oncogenic calls for hotspot variants showed good agreement across the various CDSS tools applied; however, variability existed regarding reported variants and clinical interpretation. Conclusion The pilot ring test revealed clinically relevant discrepancies between laboratories and interpreters, underscoring the need for structured external quality assessment schemes for CDSS tools in addition to the existing laboratory workflow schemes. It also highlighted several challenges related to the generation of realistic synthetic data, the design of reporting formats, the definition of ground truth, and the manual interpretation of results.
V. Nygaard, S. Zhao, D. Tamborero et al.· ESMO real world data and dig...· 0 citations
PURPOSE
To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes.
METHODS
Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar's odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139.
RESULTS
This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions).
CONCLUSION
Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.
Syed Arsalan Ahmed Naqvi, I. Riaz, Amir Saeidi et al.· JCO Clinical Cancer Informat...· 0 citations
Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning remains unclear. We evaluated agentic LLM systems for breast cancer treatment recommendation generation using 72 real clinical cases across stages I to IV and 1,147 case-specific rubrics generated through Asymmetric Information Rubric Generation (AIRG), in which the rubric generator had access to real clinical decisions unavailable to the evaluated models. Seven pipelines were compared, including single-LLM baselines, tool-augmented systems, and multi-agent architectures with fact checking and autonomous subagent spawning. The best-performing configuration, Claude Opus 4.8 with the D&C+SA pipeline, achieved a global score of 0.594 $\pm$ 0.025. Tool use and increased agent autonomy had mixed effects, improving performance in some settings but degrading it in others. Performance varied by clinical domain and disease stage, and oncologist-led error analysis revealed persistent clinically relevant failures, including incorrect or missing recommendations, flawed justifications, citation errors, outdated claims, and overconfidence. These findings suggest that agentic LLM systems can generate clinically relevant breast cancer recommendations, but remain insufficient for unsupervised clinical use.
Vinicius Anjos de Almeida, N. H. Borges, Leonardo Vicenzi et al.· 0 citations
Abstract Background Maintenance of oncology clinical practice guidelines (CPGs) is increasingly challenged by the rapid growth of trial data and therapeutic complexity. While large language models (LLMs) have shown promise in information retrieval, their utility in the rigorous, end-to-end workflow of guideline maintenance remains underexplored. Objective This case study aimed to systematically evaluate the performance of frontier LLMs in supporting oncology guideline maintenance. We sought to determine their reliability in predicting necessary guideline updates based on new evidence, their accuracy in extracting data from clinical trials, and their effectiveness as automated auditors for detecting errors in established guidelines. Methods Using the Onkopedia peripheral T-cell lymphoma (PTCL) guideline as a prospective case study, we tasked frontier models with deep-research modes and autonomous web-search capabilities (Gemini 2.5 Pro and GPT o4-mini-high) to predict a guideline update in August 2025 based on the 2021 version. Predictions were validated against the official 2025 revision published in October 2025. Next, we benchmarked evidence extraction accuracy across 80 pivotal trials using models of varying scale (27B-671B parameters vs frontier). Finally, we deployed a stacked LLM workflow to audit 28 recently updated Onkopedia guidelines for linguistic and content-related errors. Results In the predictive task, models captured 36.7% to 40% of substantive updates, often identifying landmark approvals, but frequently overstating evidence. An independent, model-blinded rescoring yielded substantial agreement (weighted Cohen κ=0.75) and confirmed predictive accuracies of 35% to 38.3%. While frontier models demonstrated high accuracy (up to 99.2%) in extracting data from individual studies, substantially outperforming smaller open-source models, this precision declined during multisource synthesis. We observed a position-dependent performance drop in long-form generation, with GPT’s endpoint accuracy dropping from 84.6% in the first half of the drafted guideline to 38.5% in the second half. As automated auditors of existing CPGs, the models successfully identified a median of 16.5 (IQR 13.8-20.3) formal errors per document and detected several clinically relevant inconsistencies (eg, invalid scoring formulas and incorrect staging definitions). Conclusions LLMs currently lack the reasoning stability for autonomous guideline authoring due to deficits in complex synthesis. However, they are effective tools for high-fidelity evidence extraction and automated quality assurance, supporting a human-led, AI-augmented workflow for efficient guideline maintenance.
M. Knauer, Julian Greß, J. Kather et al.· JMIR AI· 0 citations