RadPRISM makes a clinician-defined radiology schema a designated stratification axis: an on-premise large language model extracts per-concept text spans from free-text reports, and each clinical concept is aligned in its own dedicated visual subspace, turning concept stratification into direct, top-level alignment supervision.
Abstract
Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility. We introduce RadPRISM, which makes a clinician-defined radiology schema a designated stratification axis: an on-premise large language model extracts per-concept text spans from free-text reports, and each clinical concept is aligned in its own dedicated visual subspace, turning concept stratification into direct, top-level alignment supervision. Instantiated on chest radiographs with a 19-concept schema over $203{,}602$ examinations from an internal multi-year archive, RadPRISM improved internal dataset zero-shot classification from $0.717$ (95% CI, $0.710-0.723$) to $0.868$ (95% CI, $0.863-0.872$) macro AUROC over a matched global-alignment baseline, performed on par with the purpose-built CARZero reference in external zero-shot classification while substantially outperforming it (up to 4.3-fold) in pointing-game visual grounding. In addition, a radiologist reader study demonstrated concept-stratified retrieval ability ($0.78$ macro retrieval correctness rate within rank 3), surfacing disentangled descriptive findings that report-level retrieval and fixed-label vocabularies cannot express. RadPRISM yields discriminative, spatially faithful, natively concept-stratified representations shaped by and transparently inspectable by clinicians.
This work proposes Graph-Supervised Hierarchical Clinical Alignment, which reformulates image-report supervision as a hierarchical clinical alignment problem, and combines instance-conditioned discriminative matching with disease-conditioned soft regularization, enabling fine-grained yet clinically consistent cross-mod...
Ying-Shu Li, Yunyi Liu, Zhan-Yu Wang et al.· 0 citations
Vision-language models (VLMs) are increasingly applied to medical imaging, yet public benchmarks may reward memorization over perception: their images and questions can enter pretraining corpora, and many items remain answerable from question text alone. We present an automated, agent-driven pipeline that builds multip...
Bo Liu, Han Gu, Xiang-Rui Li et al.· Research Square· 0 citations
Vision-Language Pretrained Models (VLPMs) offer a scalable path to open-vocabulary chest radiology understanding, yet two aspects remain underexplored: how structured clinical semantics extracted from medical reports can reduce in-batch noise during contrastive learning, and how cross-modal fusion can be designed to pr...
Jianzhong You, Yuan Gao, Chris McIntosh· 0 citations
A clinically curated Pan-Asia WSI--report dataset is introduced and the REG 2025 benchmark is established as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology model...
Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al.· 0 citations
Chest X-ray report generation systems are valuable for assisting disease diagnosis and improving healthcare efficiency. However, existing methods still face two key challenges. First, multiple diseases often co-occur, leading to a combinatorial explosion of label combinations and sparse supervision for learning a gener...
Hong-Ze Zhu, Hong Liu, Ya-Wen Huang et al.· IEEE Transactions on Medical...· 0 citations
Deep learning classifiers for breast cancer histology achieve expert-level accuracy but do not explain which tissue regions drove the diagnosis. We present
GRALIS-Report
, an attribution pipeline with two defining architectural properties:
raw images never enter the report generation stage
, and
no informati...
Raimondo Fanale· Frontiers in Imaging· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.