Skip to content

CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

Aug 2026 · 0 citations · 44 references
Computer Science

Abstract

In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrence detection, and ongoing patient management. Yet, despite this central role of temporal comparison in clinical decision-making, existing medical foundation models remain largely confined to single-study understanding, leaving temporally grounded cross-examination insufficiently addressed. To address this gap, we study longitudinal imaging difference reporting, a task in which a model takes two temporally separated scans from the same patient and generates a clinically meaningful report describing interval changes between them. We introduce CT-$\Delta$Bench, a dedicated benchmark for this task with patient-level splitting to prevent information leakage. To better evaluate this task beyond surface-level text similarity, we further develop change-aware metrics specifically designed to capture clinically meaningful longitudinal changes, and conduct an independent physician validation to assess the reliability of the synthesized references and event extraction pipeline. We also compare direct paired-CT reasoning with an indirect two-stage pipeline that first generates single-timepoint reports and then performs textual differencing. Finally, we propose DeltaMed, a baseline model for direct paired-CT difference reporting, and train it on the benchmark training set. Together, these contributions lay the groundwork for temporally aware medical foundation models that better reflect real-world longitudinal clinical reasoning.

View source

Similar papers

Preprint Aug 2026

How Good are Foundation Models in Longitudinal MRI Disease Progression Reasoning?

Magnetic Resonance Imaging (MRI) interpretation is fundamental to clinical decision-making, requiring radiologists to integrate multi-view anatomical planes across sequential timepoints while precisely localizing interval changes. However, existing vision-language benchmarks remain confined to single-timepoint, single-...

Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura et al. · 0 citations
Preprint Aug 2026

ALTER: Modeling Longitudinal Changes via Regional Differencing for 3D CT Report Generation

Computed tomography (CT) is widely used for clinical diagnosis and longitudinal follow-up, yet automatically generating accurate and complete radiology reports from three-dimensional (3D) CT remains challenging. Existing methods improve fine-grained correspondence between images and text by modeling anatomical regions,...

Dong-Chen Li, Jitao Liang, Wei Li · 0 citations
Review Open access Sep 2026

Understanding and Mitigating Distribution Shifts in Volumetric Lung Nodule CAD Using a 3D Vision-Language Framework.

As artificial intelligence becomes increasingly integrated into medical imaging practice, its robustness across heterogeneous real-world settings remains a major challenge. We quantified the effect of real-world distribution shifts on three-dimensional AI models for lung nodule analysis on CT and, motivated by these sh...

B. Bercean, Rafael Medelean, A. Tenescu et al. · 0 citations
Preprint Aug 2026

Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation

An architecture-agnostic framework is proposed that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model that improves anatomical sensitivity through explicit visual grounding and provides an independent interpretability output for clinical oversight, mo...

Bruno Palau, Franziska Vogt, Daria Laslo et al. · 0 citations
Preprint Aug 2026

Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models

Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models could assist this process by providing generalizable representations of anatomy and pathology. To evaluate their diagnostic breadth, we benchmark ten frozen CT encoders across thre...

Maulik Chevli, Johannes Brandt, R. Braren et al. · 0 citations
Preprint Sep 2026

A multicenter benchmark and clinically structured metric for coronary CTA report generation

Reliable evaluation of automated coronary computed tomography angiography (CCTA) report generation requires standardized multicentre benchmarks and clinically structured metrics. We established a four-centre benchmark comprising 3,021 CCTA series from 818 patient-report pairs to evaluate seven open-source three-dimensi...

Zhi-Yu Ye, Yue Sun, Li-Miao Zou et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.