Skip to content
Open access

Explainable Multi-Omics Transformer Risk Stratification in Lung Adenocarcinoma: External Validation, Calibration, and Patient-Level Genomic Attribution

Aug 2026 · Iconic research and engineering journals · 0 citations · 23 references

Abstract

- Lung adenocarcinoma (LUAD) remains biologically heterogeneous and clinically difficult to stratify with a single data modality. Transformer architectures can, in principle, model cross-modality interactions among RNA expression, somatic mutation, copy-number variation (CNV), and clinical variables, but their clinical value depends on reproducible validation, calibration, and patient-level explainability rather than architectural novelty alone. This article develops a revised validation and reporting framework for an explainable multi-omics transformer risk-stratification model in LUAD. The study aligns TCGA-LUAD development results, complete-case multi-omics survival modelling, and MSK-IMPACT external validation evidence to test whether transformer fusion adds clinically meaningful, transportable, and interpretable patient-level information while preserving acceptable discrimination and calibration. The empirical evidence shows a nuanced pattern. A RNA-plus-clinical ridge Cox baseline achieved the strongest apparent discrimination (C-index = 0.724), while the full multi-omics ridge Cox model produced lower aggregate discrimination (C-index = 0.682) but stronger biological interpretability, temporal stability, improved calibration logic, and positive decision-curve value. The transformer models produced moderate internal discrimination (RNA-only C-index = 0.6333; full multi-omics transformer C-index = 0.6094) and modest but non-trivial external transportability on MSK-IMPACT LUAD (C-index = 0.5368), while retaining patient-level mechanistic attribution through Integrated Gradients. Attribution analysis indicated that RNA expression dominated model explanations, contributing approximately 70% of attribution mass, followed by somatic mutations (15%), CNVs (10%), and clinical variables (5%). Subtype-level interpretation was biologically coherent: high-risk cases showed TP53 activity, cell-cycle dysregulation and DNA repair defects; intermediate-risk cases showed metabolic reprogramming; and low-risk cases showed immune-rich signatures. The paper argues that the strongest claim for explainable transformer-based multi-omics LUAD modelling is not simple predictive superiority over classical baselines. Rather, its defensible contribution is a validated, auditable, and patient-level framework that can connect discrimination, calibration, transportability, and biologically intelligible explanation in a form suitable for further prospective evaluation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.