This work proposes Graph-Supervised Hierarchical Clinical Alignment, which reformulates image-report supervision as a hierarchical clinical alignment problem, and combines instance-conditioned discriminative matching with disease-conditioned soft regularization, enabling fine-grained yet clinically consistent cross-modal representations.
Abstract
Radiology report generation (RRG) has recently benefited from large language models, which substantially improve report fluency. However, clinically faithful generation remains challenging because current supervision is still imposed mostly at the report level. This creates a granularity mismatch: radiology reports are composed of disease-grounded findings, while existing methods are trained mainly with whole-report objectives. To address this problem, we propose Graph-Supervised Hierarchical Clinical Alignment, which reformulates image-report supervision as a hierarchical clinical alignment problem. Our method structures this alignment as a disease-conditioned process, where supervision is decomposed into two levels: Disease-Centric Alignment for fine-grained disease-specific correspondence, and Global Clinical Semantic Alignment for report-level semantic coherence. A clinical knowledge graph is used as a training-time-only structural prior that defines disease-specific supervision units and their clinical relationships, introducing no additional overhead at inference. Because standard contrastive alignment could produce false negatives when studies share overlapping pathologies, we combine instance-conditioned discriminative matching with disease-conditioned soft regularization, enabling fine-grained yet clinically consistent cross-modal representations. Experiments on MIMIC-CXR, IU-Xray, and COV-CTR show that our method consistently improves performance on both conventional and clinical metrics. Notably, our 3B model surpasses several prior systems with larger 7B/13B backbones, suggesting that improving supervision structure, rather than increasing model size, can be more effective for RRG.
A systematic study spanning five KG task formulations, three training paradigms, two KGs, and three base LLMs finds that at the task level, all paradigms improve over the non-finetuned baseline, but methods with comparable in-domain accuracy show substantially different knowledge transfer behavior.
Saksham Khatwani, He Cheng, M. Afshar et al.· 0 citations
ClinAlign—a memory-based retrieval framework aligned with clinical workflow, drawing inspiration from clinical diagnostic workflows is proposed, which constructs a disease-aware visual memory bank and introduces Classification-Guided Prompt Augmentation (CGPA), where disease state predictions are converted into structu...
Lihong Qiao, Shi-Yi Gao, Yu-Cheng Shu et al.· Proceedings of the Thirty-Fi...· 0 citations
A clinically curated Pan-Asia WSI--report dataset is introduced and the REG 2025 benchmark is established as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology model...
Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al.· 0 citations
The proposed AG-VLM framework provides a scalable foundation for computer-assisted radiology reporting while retaining the need for radiologist verification before clinical use and indicates that explicit attention-guided visual reasoning combined with cross-modal semantic alignment can generate more accurate, clinical...
P. Dayaker, M. Vignesh, I. Z. et al.· International journal of com...· 0 citations
Experiments on BUS-CoT and IU X-ray datasets demonstrate consistent improvements in diagnostic accuracy, concept consistency, and report quality over strong general-purpose and medical MLLMs, indicating that concept-grounded reasoning better aligns generation with clinical decision processes.
Xin-Yue Xu, Hong-Bin Lin, Juan-Gui Xu et al.· 0 citations
Chest X-ray report generation systems are valuable for assisting disease diagnosis and improving healthcare efficiency. However, existing methods still face two key challenges. First, multiple diseases often co-occur, leading to a combinatorial explosion of label combinations and sparse supervision for learning a gener...
Hong-Ze Zhu, Hong Liu, Ya-Wen Huang et al.· IEEE Transactions on Medical...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.