Skip to content
Preprint

From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

Aug 2026 · 0 citations · 112 references
Computer Science

TL;DR

Multi-Resolution Pyramid Transformer (MRPT) is introduced, a model that hierarchically aggregates multi-resolution information from cellular to tissue and WSI levels and surpasses recent foundation models and Multimodal Large Language Models in cancer subtype classification, tissue phenotyping, and Visual Question Answering for WSI understanding.

Abstract

Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution Whole Slide Images (WSIs), limiting their generalization across arbitrary resolutions. Gigapixel WSIs inherently contain diagnostic patterns at multiple scales, including cellular morphologies, tissue architectures, and global context, mirroring how expert pathologists examine WSIs. We introduce Multi-Resolution Pyramid Transformer (MRPT), a model that hierarchically aggregates multi-resolution information from cellular to tissue and WSI levels. MRPT employs a biologically meaningful Consecutive Cross-Resolution Attention (CCRA) mechanism to capture scale-independent interactions and enforces multi-resolution semantic consistency by aligning embeddings across resolutions, yielding robust and generalizable WSI representations. Pre-trained in a multi-resolution self-supervised manner on 624M patches, 2.4M regions, and 36K WSIs, MRPT learns rich coarse-to-fine histopathology features. Extensive experiments on 34 diverse datasets show that MRPT surpasses recent foundation models and Multimodal Large Language Models (MLLMs) in cancer subtype classification, tissue phenotyping, and Visual Question Answering (VQA) for WSI understanding.

View source

Similar papers

Preprint Sep 2026

Learning Where to Focus: Self-Supervised Multi-Scale ViTs for Histopathology

Pathologists diagnose diseases by first locating suspicious tissue and then examining it at higher magnification, whereas self-supervised vision transformers (ViTs) allocate the same spatial resolution to every image region despite diagnostic evidence being sparse and spanning multiple biological scales. Recent patholo...

Anabel Stammer, Valay Bundele, Mehran Hosseinzadeh et al. · 0 citations
#machine learning Preprint Sep 2026

Nonparametric Distribution Matching for Self-Supervised Whole-Slide Image Condensation

Histological whole-slide images (WSIs) are central to computational pathology but pose severe computational challenges due to their extremely high resolution, often spanning several gigabytes per slide. To enable scalable learning, existing methods apply self-supervised data condensation to reduce computational cost, b...

Duong M. Nguyen, T. Hoang, H. Nguyen et al. · 1 citation
Preprint Aug 2026

One Model to Magnify Them All: Efficient Scale-Invariant Histopathology via Conditional Normalization and Continuous Magnification Training

Whole slide images (WSIs) in digital histopathology are acquired at discrete magnification levels encoding complementary diagnostic information from global tissue architecture to fine-grained cellular morphology. Yet, deep learning models remain sensitive to scale variation. Existing magnification-invariant methods rel...

Agnieszka Florkowska, H. Müller, M. Wodzinski · 0 citations
Preprint Sep 2026

Towards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images

Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell reso...

Zi-Jun Gao, Chun-Bin Gu, Jin-Xi Xiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.