Skip to content
Open access

Hierarchical vision transformers for Epstein-Barr virus status and histological subtype prediction in Hodgkin lymphoma whole-slide images

Jul 2026 · Journal of Pathology Informatics · Vol 22 · 0 citations · 56 references
Medicine

Abstract

Accurate stratification of Hodgkin lymphoma (HL) by immunologic/histological subtypes and Epstein-Barr virus (EBV) status is essential for epidemiological and translational research, yet large-scale testing is impractical and expensive. Digital pathology models that utilize routinely used hematoxylin and eosin (H&E) whole-slide images (WSIs) could close this gap. We developed and validated a hierarchical Vision Transformer pipeline that aggregates cell-, patch-, and region-level context to predict EBV status and the three most prevalent immunological/histological HL subtypes: nodular sclerosis (NS), mixed cellularity (MC), and nodular lymphocyte-predominant HL (NLPHL)—from H&E-stained WSIs, and additionally evaluated a standard attention-based multiple-instance learning (ABMIL) baseline for direct architectural comparison. The development pool comprised 1643 HL cases (1952 WSIs) from 18 Danish hospitals and was used for hospital-preserving 5-fold cross-validation; external validation was performed on an independent hold-out cohort of 458 cases (532 WSIs) from five hold-out hospitals. For subtype prediction, analyses were restricted to the 1560 cases belonging to NS, MC, or NLPHL. On the external EBV cohort (N=458) the hierarchical pipeline achieved an area under the receiver operating characteristic curve (ROC–AUC) of 0.73 (95% confidence interval (CI) 0.68–0.77), precision–recall (PR)–AUC 0.57 (95% CI 0.49–0.66) with recall (sensitivity) 0.74 (95% CI 0.67–0.80) and macro-F1 score 0.60 (95% CI 0.54, 0.65). For 3-class subtype prediction on 359 external cases, discrimination reached ROC–AUC 0.84 (95% CI 0.80–0.88) and PR–AUC 0.63 (95% CI 0.56–0.71) with a macro-F1 of 0.56 (95% CI 0.48, 0.64) and macro-recall 0.55 (95% CI 0.47, 0.63); residual errors were dominated by NS-MC confusions. An ABMIL baseline using the same patch embeddings achieved ROC–AUC 0.76 (95% CI 0.72–0.81) for EBV and 0.89 (95% CI 0.86–0.92) subtype prediction, outperforming the hierarchical model on both tasks. This multicenter study shows that both hierarchical and attention-based architectures can determine EBV status and major HL subtypes directly from routine H&E slides with externally validated performance across hospitals, whereas the finding that the simpler baseline outperformed the hierarchical model suggests that strong foundation-model embeddings combined with attention-based pooling may reduce the need for explicit multi-scale modelling in cohorts of this size.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.