Skip to content

Solving the Needle-in-a-Haystack Problem in Mammography Vision-Language Model with Differentiable Subset Sampling

Sep 2026 · 0 citations · 82 references
Computer Science

TL;DR

TopKSigLIP outperforms existing open-source mammography and general medical VLMs on both internal and external benchmarks on density assessment, BI-RADS classification, finding subtyping, and cancer prediction under zero-shot evaluation.

Abstract

There is growing interest in adopting CLIP-style vision--language model (VLM) pretraining for mammography. However, models that directly employ the standard CLIP architecture and training objective exhibit limited zero-shot performance in clinically important tasks such as cancer, finding-type, and BI-RADS predictions. We argue that this underwhelming performance is due to neglecting two characteristics of mammography data: (1) its high-res nature, and (2) homogeneity of radiology reports, largely driven by a predominance of negative/benign findings on examinations. We propose TopKSigLIP, a VLM designed to address these two limitations through a novel architecture and learning objectives. Instead of downscaling high-res mammography images to satisfy GPU memory constraints, TopKSigLIP introduces TopK-Patch module that learns to sample a sparse set of high-res patches likely to contain lesions, sidestepping the resolution--batch size tradeoff of VLM training. The sampled patch locations additionally serve as a built-in localization tool. To address report homogeneity, we replace the contrastive loss, which falsely repels semantically similar pairs, with a Sup-sigmoid loss. Sup-sigmoid loss extends the sigmoid loss from SigLIP with soft labels derived from structured data. TopKSigLIP outperforms existing open-source mammography and general medical VLMs on both internal and external benchmarks on density assessment, BI-RADS classification, finding subtyping, and cancer prediction under zero-shot evaluation. TopKSigLIP remains competitive under linear probing despite using a significantly smaller vision encoder and smaller training batches than baselines. The TopK-Patch module additionally achieves superior lesion localization over post-hoc Grad-CAM. Code and weights are made public:https://github.com/Youngseok0001/TopKSigLIP.

View source

Similar papers

Open access Sep 2026

LENS: A mammography-specific hybrid CNN-Transformer with lesion-aware evidence modeling

Background Artificial Intelligence (AI) models for mammography classification is prone to shortcut learning because diagnostically relevant evidence is typically sparse, localized, and easily dominated by non-lesion background context. This study aimed to develop a mammography-specific framework that integrates lesion-...

Duc Quy Hoang, Van Kien Cao, Tan-Nhu Nguyen et al. · 0 citations
Preprint Sep 2026

Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images

Automating the analysis of whole-slide images (WSIs), a key step in cancer diagnosis, has high clinical value, as it can reduce pathologist's workload while improving diagnosis accuracy. Recently, vision-language models have shown promising performance for patch-level classification without requiring any annotation, ye...

T. Godelaine, M. Dausort, Karim El Khoury et al. · 0 citations
Preprint Aug 2026

DualMiT-Net: Local-Global Transformer-Convolutional Fusion for Breast Mass Segmentation in Mammographic Regions of Interest

Breast mass segmentation is an important step in computer-aided mammography, but it remains difficult because masses can have low contrast, irregular shapes, and boundaries that blend with surrounding breast tissue. To address this problem, we present DualMiT-Net, a dual-branch network that uses both a focused view of...

Alibek Kamiluly, M. Muratova, Yash Patel et al. · 0 citations
Open access Sep 2026

BCSeg-IRUNet: A Hybrid Inception–Residual Encoder–Decoder for Mammographic Mass Segmentation

Breast cancer accounted for approximately 2.4 million new cases and 694,000 deaths worldwide in 2024. Deep learning approaches, particularly encoder–decoder architectures, have been investigated for mammographic mass segmentation. However, evaluations are commonly centered on benchmarked datasets, which include only an...

Fabian Cienfuegos-Caraveo, A. Guzmán-Pando, G. Ramírez-Alonso et al. · 0 citations
Preprint Sep 2026

Weakly-supervised Kidney Tumor Classification from CT Scans with Multi-Instance Learning and Anatomical Filtering

Deep learning models for CT scan analysis are often limited by the scarcity of precise pixel-level annotations, which require significant radiologist effort to produce. Training on scan-level labels alone reduces annotation requirements but introduces challenges: low supervision ratios and large input volumes make mode...

Joonas Ariva, Dmytro Fishman · 0 citations
#artificial intelligence Preprint Sep 2026

ProtoCAM: Interpretable Few-Shot Mask-Guided Prototypical Learning for Breast Lesion Classification in Ultrasound Imaging

Breast ultrasound imaging plays an important role in the early detection and diagnosis of breast cancer, particularly for patients with dense breast tissue. However, developing reliable deep learning models for ultrasound analysis is challenging due to limited annotated medical data and the need for interpretable predi...

Ashkan Ebadi · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.