Skip to content
Conference

Post-Hoc Explainability for Generative Volumetric Medical Vision-Language Models

Jul 2026 · 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET) · pp. 1-6 · 0 citations · 18 references

Abstract

In radiology, multimodal vision-language models (VLMs) are increasingly used for clinical tasks such as report generation and visual question answering. Their adoption has raised important concerns regarding the transparency and trust-worthiness of the generated text reports, as the reasoning process leading to a given report remains opaque for clinicians. In this work, we investigate how post-hoc attribution methods behave in large generative volumetric VLMs that jointly process 3D scans and clinical text, a setting that remains largely unexplored. We propose to use Post-hoc gradient-based attribution that directly links small changes in the input volume to changes in the model's output probability. The faithfulness and spatial specificity of four attribution methods are subsequently characterized and compared when applied to Med3DVLM, a recent medical VLM for image-text understanding. To enable differentiable attribution at inference time, we apply teacher forcing exclusively during the attribution forward pass so that the stochastic generation path is converted into a differentiable computation graph. The cumulative log-likelihood of the generated response is used as the scalar attribution target. Input-level methods, including Saliency and Integrated Gradients, alongside feature-level techniques such as Grad-CAM and Guided Grad-CAM, are computed on volumetric inputs across diverse clinical question types. Qualitative assessment and a quantitative voxel deletion protocol indicate that input-space gradient methods, especially Integrated Gradients, produce spatially selective and causally faithful relevance maps, whereas feature-level methods like Layer Grad-CAM exhibit significant spatial diffusion and fail to provide discriminative localization in multimodal architectures. These results demonstrate that input-level attributions exhibit a significantly steeper drop in model confidence compared to feature-level methods, confirming they are more faithful and spatially precise, while intermediate-layer maps remain diffuse and non-discriminative.

View source