Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

PanoraOnc: A pan-cancer clinico-genomic AI model for transferable outcome predictions

Progress in precision oncology, including biomarker discovery and individualized treatment selection, is limited by the complexity of clinico-genomic data and the scarcity of large multimodal patient cohorts. Here, we introduce PanoraOnc, a pan-cancer artificial intelligence (AI) model pretrained on real-world clinical, genomic, and imaging data from 84,131 patients spanning 66 cancer types. PanoraOnc enables transferable treatment outcome prediction through pan-cancer pretraining and generalizes to unseen cohorts across cancer types, institutions, and therapeutic settings. Evaluation and fine-tuning were performed on cohorts comprising diverse modalities, including clinical features, targeted gene panels, immunofluorescence imaging, whole-exome sequencing, and transcriptomic profiles. Across these settings, PanoraOnc consistently outperforms statistical, machine-learning, survival, and AI baselines, with the largest improvements observed in zero- and few-shot scenarios, demonstrating that large-scale clinico-genomic pretraining enables robust and generalizable outcome predictions across previously unseen conditions. In addition, PanoraOnc supports biomarker discovery through explainable AI, revealing both established and underappreciated features, including tumor-infiltrating clonal hematopoiesis, oncogenic signaling pathways, and DNA damage response mechanisms in immunotherapy-treated melanoma and non-small cell lung cancer. Furthermore, PanoraOnc enables the identification of patient subgroups potentially benefitting from alternative treatments by estimating personalized treatment outcomes across therapeutic scenarios. These findings establish pan-cancer multimodal pretraining as a scalable paradigm for AI-assisted discovery in precision oncology.

M. Schuerch, J. Geisberg, C. T. Flower et al. · 0 citations
Jul 2026

Strategies for Deploying Large Language Models for Ascertaining Clinical Outcomes and Sites of Metastases From Radiology Impressions in Patients With Cancer.

PURPOSE To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes. METHODS Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar's odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139. RESULTS This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions). CONCLUSION Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.

Syed Arsalan Ahmed Naqvi, I. Riaz, Amir Saeidi et al. · 0 citations