Skip to content
Review

Abstract A025: Real-World Data–Driven Target Discovery in Advanced NSCLC

Jul 2026 · Clinical Cancer Research · Vol 32, pp. A025-A025 · 0 citations

TL;DR

This work demonstrates the novel application and value of multimodal liquid-biopsy RWD for scalable, clinically anchored therapeutic Target ID in solid tumors and indicates a complementary discovery paradigm that accelerates hypothesis generation grounded in patient outcomes.

Abstract

Real-world data (RWD) is increasingly leveraged to inform clinical development and health-outcomes research in oncology; however, its use in therapeutic target identification (Target ID) remains underexplored. Emerging non-invasive liquid biopsy technologies now enable large-scale, longitudinal molecular profiling within routine care, offering a unique opportunity to derive biologically rich, clinically grounded insights at population scale. We assessed whether multimodal RWD—including blood-based genomic and epigenomic NGS—combined with machine learning (ML) could generate actionable Target ID hypotheses in advanced non–small cell lung cancer (NSCLC). We analyzed a de-identified cohort of several hundred patients with advanced NSCLC from Guardant's InfinityAI Data Library, a multimodal oncology real-world data platform integrating longitudinal clinical records, treatment histories, outcomes, and liquid biopsy–derived genomic and epigenomic profiles. The analysis followed a two-phase framework. Phase 1 used baseline, pre-treatment liquid biopsy data to identify molecular features and pathway-level programs associated with primary resistance, defined using real-world response metrics including time-to-discontinuation and time-to-next-treatment. Phase 2 extends this framework to longitudinal post-treatment profiling to characterize treatment-induced molecular changes and emergent dependencies. Machine learning–derived candidates are prioritized using literature review, ClinicalTrials.gov, patent analysis, and public functional datasets (TCGA, DepMap). ML analysis revealed reproducible molecular signals and pathways associated with divergent clinical outcomes, including features extending beyond canonical oncogenic drivers. Incorporation of epigenomic RWD substantially improved predictive performance and provided mechanistic insight not captured by genomic alterations alone. The prioritized targets recapitulated known resistance mechanisms while uncovering previously unrecognized, potentially actionable pathways implicated in treatment resistance and disease progression. This work demonstrates the novel application and value of multimodal liquid-biopsy RWD for scalable, clinically anchored therapeutic Target ID in solid tumors. Integrating real-world molecular and clinical data with ML offers a complementary discovery paradigm that accelerates hypothesis generation grounded in patient outcomes. Ongoing in vitro studies across diverse cancer cell lines will further interrogate prioritized targets and support translation toward therapeutic development. Aaron Hardin, Peili Zhang, Amar Das. Real-World Data–Driven Target Discovery in Advanced NSCLC [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr A025.

View source

Similar papers

Open access Aug 2026

AI-HOPE lung cancer: a multicenter real-world registry integrating artificial intelligence for metastatic non-small-cell lung cancer

Background The AI-HOPE Lung Cancer study is a multicenter initiative designed to integrate artificial intelligence (AI) and real-world data to improve outcome prediction in patients with metastatic non-small-cell lung cancer treated with first-line immunotherapy-based regimens. AI-HOPE aims to leverage machine learning (ML) models to generate individualized predictions of progression-free survival (PFS), overall survival (OS), and treatment-related toxicity in a broad, unselected population. Materials and methods Clinical and imaging data are harmonized and stored within a privacy-compliant infrastructure (San Raffaele Ai CEnter [S-RACE] platform), promoting FAIR (Findable, Accessible, Interoperable and Reusable) data principles and minimizing manual workload. The primary objective is the development of time-to-event models for PFS and OS. Complementary binary classification models will explore early progression, long-term survival, and clinically relevant toxicities. Results The study includes retrospective (from 2017) and prospective (until 2027) phases across 21 European centers. So far, 920 patients have been recruited for the study, of whom 621 have baseline imaging scans available for centralized analysis. In the AI-HOPE study, a flexible methodological approach integrates multiple ML models tailored to specific clinical questions, complemented by explainable AI tools. Multimodal models combining clinical variables with computed tomography and [18F]2-fluoro-2-deoxy-d-glucose–positron emission tomography imaging features (when available) are supported through the S-RACE platform, which provides a partially automated imaging analysis workflow. Conclusions By combining structured clinical variables and multimodal imaging data, the AI-HOPE Lung Cancer study aims to support refined risk stratification and treatment personalization, ultimately facilitating the responsible integration of AI into routine thoracic oncology practice.

F. Ogliari, M. Ferrara, J. Huijs et al. · 0 citations
Open access Aug 2026

PanoraOnc: A pan-cancer clinico-genomic AI model for transferable outcome predictions

Progress in precision oncology, including biomarker discovery and individualized treatment selection, is limited by the complexity of clinico-genomic data and the scarcity of large multimodal patient cohorts. Here, we introduce PanoraOnc, a pan-cancer artificial intelligence (AI) model pretrained on real-world clinical, genomic, and imaging data from 84,131 patients spanning 66 cancer types. PanoraOnc enables transferable treatment outcome prediction through pan-cancer pretraining and generalizes to unseen cohorts across cancer types, institutions, and therapeutic settings. Evaluation and fine-tuning were performed on cohorts comprising diverse modalities, including clinical features, targeted gene panels, immunofluorescence imaging, whole-exome sequencing, and transcriptomic profiles. Across these settings, PanoraOnc consistently outperforms statistical, machine-learning, survival, and AI baselines, with the largest improvements observed in zero- and few-shot scenarios, demonstrating that large-scale clinico-genomic pretraining enables robust and generalizable outcome predictions across previously unseen conditions. In addition, PanoraOnc supports biomarker discovery through explainable AI, revealing both established and underappreciated features, including tumor-infiltrating clonal hematopoiesis, oncogenic signaling pathways, and DNA damage response mechanisms in immunotherapy-treated melanoma and non-small cell lung cancer. Furthermore, PanoraOnc enables the identification of patient subgroups potentially benefitting from alternative treatments by estimating personalized treatment outcomes across therapeutic scenarios. These findings establish pan-cancer multimodal pretraining as a scalable paradigm for AI-assisted discovery in precision oncology.

M. Schuerch, J. Geisberg, C. T. Flower et al. · 0 citations

Explainable AI for analyzing cancer outcomes using large-scale genome sequencing data

Metastatic cancer remains a leading cause of global mortality, yet accurate prognosis is frequently hampered by high-dimensional molecular features and heterogeneous clinical presentations. While traditional staging systems and linear models provide a foundational risk assessment, they often fail to capture the complex, nonlinear interactions between metastatic topology, genomic burden, and functional sequence variation. To address this, recent advances in machine learning and genomic foundation models present a transformative opportunity to integrate diverse data types into an explainable predictive framework. Consequently, this research developed a multi-tier, explainable AI framework designed to risk-stratify patients and predict overall survival using clinical and genomic covariates. Additionally, the framework aimed to surface sequence-level disease drivers by implementing joint variant calling from RNA-seq data and leveraging transformer-based architectures. The study employed a two-track methodological approach encompassing populationscale modeling and sequence-level deep learning. For the population-scale aim, a retrospective analysis was conducted on the Memorial Sloan Kettering-Metastatic cohort, consisting of 25,775 patients. Five distinct classifiers XGBoost, Logistic Regression, Random Forest, Decision Tree, and Naive Bayes were trained on a balanced subset of 20,338 patients utilizing an 80/20 stratified split. Model explainability was established through Shapley Additive Explanations (SHAP), while survival dynamics were evaluated using Kaplan-Meier estimates, Cox proportional hazards models, and an XGBoost-Cox variant. Concurrently, a pilot study involving 60 individuals, comprising 30 breast cancer cases and 30 controls, investigated sequence-level drivers using RNAseq data. A joint variant calling pipeline generated a unified genomic variant call format for association testing, and three genomic foundation models DNABERT-2, HyenaDNA, and Nucleotide Transformer were fine-tuned for 50 epochs on variantcentered windows spanning 100 base pairs in either direction to classify case versus control status. The results revealed stark contrasts in performance between the clinical and genomic modeling tracks. In survivability predictions, XGBoost emerged as the superior classifier, achieving an accuracy of 0.74 and an AUC of 0.82, while the XGBoost-Cox model outperformed the traditional Cox model with a C-index of 0.70 compared to 0.66. Through explainability and hazard-based analyses, metastatic site count, tumor mutational burden, the fraction of the genome altered, and the presence of liver and bone metastases were identified as the most potent prognostic indicators across pan-cancer and cancer-specific models. Conversely, the sequence-level transformer models exhibited severe overfitting, with test performance remaining near stochastic levels between 49 percent and 51 percent accuracy. Although DNABERT-2 achieved the highest nominal accuracy at 50.63 percent and HyenaDNA showed superior computational efficiency, the pilot ultimately indicated that fine-tuning transformers on raw sequences in small cohorts is heavily limited by a high signal-to-noise ratio and the polygenic complexity of cancer. Ultimately, this research demonstrates that explainable machine learning models can robustly predict survivability and highlight actionable features for oncology dashboards. However, future sequence-level deep learning efforts must pivot toward using frozen transformer embEd. D.ings or larger, multi-center cohorts to ensure equitable and generalizable clinical adoption.

P. Nalela · 0 citations
Open access Jul 2026

Advancing cancer detection and treatment using longitudinal routine clinical data.

Cancer management remains fragmented across its continuum, from late-stage diagnosis and salvage therapies to non-personalized surveillance. Here, we present Oncoformer, a unified multimodal transformer model trained on the China Oncology Multimodal Prediction and Surveillance Study (COMPASS) cohort (3.67 million individuals, 17.7 million clinical visits) and validated on independent external cohorts, including the UK Biobank. Oncoformer integrates longitudinal electronic health records with chest X-ray imaging to address multiple clinical tasks: pan-cancer diagnosis (area under the receiver operating characteristic curve [AUROC] = 0.956), future cancer prediction up to 1 year before diagnosis (AUROC = 0.869), tumor stage inference (mean AUROC > 0.90), patient-specific treatment-response forecasting, and recurrence-free survival stratification across ten cancer types (all p < 0.01). Staging predictions were independently validated against postoperative pathological endpoints and shown to converge on core cancer genomic pathways. By translating routine clinical data into a dynamic view of cancer evolution, Oncoformer provides a framework for risk-informed cancer prediction and treatment stratification using routine clinical data.

Fei Liu, Kai Wang, Hui Xu et al. · 1 citation
Jul 2026

Abstract PR006: Patient-level prediction of trial outcomes with a calibrated pan-cancer foundation model

As precision oncology drives development toward narrower biomarker-defined populations, patient-level outcome prediction models can improve trial planning and decision-making. Methods that integrate both real world data (RWD) and recent trial evidence can produce better-calibrated predictions and enable clinical applications such as synthetic control arms, comparative effectiveness, and trial design optimization. We developed a calibrated pan-cancer foundation model integrating patient-level RWD with summary-level clinical trial outcomes. The model uses a transformer-based architecture to capture joint distributions across thousands of clinical and genomic features. It was trained on over 300,000 tumor biopsy records with linked clinical data, primarily from RWD sources. The model generates synthetic patient-level cohorts conditional on user-specified I/E criteria and predicts outcomes under specified treatments. An information-geometric calibration procedure aligns these predictions with published trial baseline characteristics and outcome landmarks. We validated the model across three Phase III settings: (1) To evaluate subgroup-level prediction from population-level calibration, we generated a synthetic cohort matching mNSCLC trial POSEIDON baseline characteristics and calibrated to its published control-arm OS. We then predicted OS across PD-L1 strata, histology, and KEAP1/STK11/KRAS mutation status and recapitulated published results: 90.9% (40/44) of median and 2 to 5-year OS estimates fell within 95% CIs, with median absolute deviation 2.7%. (2) To assess out-of-sample prediction in a genetic subgroup, we simulated a BRAF V600E mCRC cohort matching BREAKWATER baseline characteristics. The model was calibrated on prior unselected mCRC trials (XELOX, TRIBE) and applied without calibration to BREAKWATER or any BRAF-selected trial. Despite only five training-data patients meeting BREAKWATER I/E criteria, model-predicted OS matched observed values at 6, 12, and 18 months. (3) To demonstrate indirect head-to-head comparison without a randomized trial, we compared nab-paclitaxel plus gemcitabine (NG) and FOLFIRINOX in mPDAC. These regimens were evaluated in MPACT and PRODIGE4 respectively, with differing populations. We simulated the MPACT arm, then used entropy balancing to conform baseline characteristics to PRODIGE4, estimating NG outcomes in a healthier PRODIGE4-like population. This decomposed the published 80-day median OS gap: ∼14% was attributable to baseline demographics, with a residual 73-day FOLFIRINOX advantage. We present a framework that calibrates patient-level OS predictions to published clinical trial evidence. Pan-cancer pretraining enables transfer learning across data sources, improving prediction in narrow populations with sparse data. The model can be further fine-tuned on data to support tailored predictions across biomarkers, indications, and treatments. Together, these capabilities provide a data-efficient approach to generating patient-level evidence for clinical applications. Daniele Bertolini, Franklin Fuller, Jason Christopher, Jonathan Walsh, Samantha I . Liang, Aaron Smith. Patient-level prediction of trial outcomes with a calibrated pan-cancer foundation model [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr PR006.

D. Bertolini, F. Fuller, J. Christopher et al. · 0 citations
Jul 2026

Abstract A030: Patient-level prediction of trial outcomes with a calibrated pan-cancer foundation model

As precision oncology drives development toward narrower biomarker-defined populations, patient-level outcome prediction models can improve trial planning and decision-making. Methods that integrate both real world data (RWD) and recent trial evidence can produce better-calibrated predictions and enable clinical applications such as synthetic control arms, comparative effectiveness, and trial design optimization. We developed a calibrated pan-cancer foundation model integrating patient-level RWD with summary-level clinical trial outcomes. The model uses a transformer-based architecture to capture joint distributions across thousands of clinical and genomic features. It was trained on over 300,000 tumor biopsy records with linked clinical data, primarily from RWD sources. The model generates synthetic patient-level cohorts conditional on user-specified I/E criteria and predicts outcomes under specified treatments. An information-geometric calibration procedure aligns these predictions with published trial baseline characteristics and outcome landmarks. We validated the model across three Phase III settings: (1) To evaluate subgroup-level prediction from population-level calibration, we generated a synthetic cohort matching mNSCLC trial POSEIDON baseline characteristics and calibrated to its published control-arm OS. We then predicted OS across PD-L1 strata, histology, and KEAP1/STK11/KRAS mutation status and recapitulated published results: 90.9% (40/44) of median and 2 to 5-year OS estimates fell within 95% CIs, with median absolute deviation 2.7%. (2) To assess out-of-sample prediction in a genetic subgroup, we simulated a BRAF V600E mCRC cohort matching BREAKWATER baseline characteristics. The model was calibrated on prior unselected mCRC trials (XELOX, TRIBE) and applied without calibration to BREAKWATER or any BRAF-selected trial. Despite only five training-data patients meeting BREAKWATER I/E criteria, model-predicted OS matched observed values at 6, 12, and 18 months. (3) To demonstrate indirect head-to-head comparison without a randomized trial, we compared nab-paclitaxel plus gemcitabine (NG) and FOLFIRINOX in mPDAC. These regimens were evaluated in MPACT and PRODIGE4 respectively, with differing populations. We simulated the MPACT arm, then used entropy balancing to conform baseline characteristics to PRODIGE4, estimating NG outcomes in a healthier PRODIGE4-like population. This decomposed the published 80-day median OS gap: ∼14% was attributable to baseline demographics, with a residual 73-day FOLFIRINOX advantage. We present a framework that calibrates patient-level OS predictions to published clinical trial evidence. Pan-cancer pretraining enables transfer learning across data sources, improving prediction in narrow populations with sparse data. The model can be further fine-tuned on data to support tailored predictions across biomarkers, indications, and treatments. Together, these capabilities provide a data-efficient approach to generating patient-level evidence for clinical applications. Daniele Bertolini, Franklin Fuller, Jason Christopher, Jonathan Walsh, Samantha I . Liang, Aaron Smith. Patient-level prediction of trial outcomes with a calibrated pan-cancer foundation model [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr A030.

D. Bertolini, F. Fuller, J. Christopher et al. · 0 citations