Skip to content

Deep learning-based diagnostic report generation for low-resolution functional medical images via cross-modal visual and textual alignment

Sep 2026 · Applied intelligence (Boston) · Vol 56 · 0 citations · 59 references

TL;DR

A unified framework for automatic report generation from SPECT bone scintigrams that integrates domain-adaptive representation learning, fine-grained image–text alignment, and anatomy-guided supervision is proposed, offering a valuable pathway to achieving trustworthy and intelligent diagnostic support within nuclear medicine.

View source

Similar papers

Open access Sep 2026

Dual-branch cross-modal architecture with global-to-local feedback for radiographic image–text retrieval in chest X-rays

Radiology reports are vital for accurate diagnosis and treatment planning, yet their manual generation is time-consuming and dependent on radiologist expertise, leading to delays and inconsistent clinical decisions. Medical image–text retrieval offers a scalable solution by enabling the retrieval of relevant prior case...

Rezaul Abedin, Sofiane Laridi, Kam-Ming Mark Tam · 0 citations
Sep 2026

Concept-Enhanced Multi-Scale Cross-Modal Alignment for Medical Visual Representation Learning.

A Concept Clause Decomposition method is designed to extract semantically complete descriptions of pathological findings or radiology manifestations from medical reports as medical concept clauses, which are then utilized within a multi-granularity cross-modal alignment framework to enhance medical concept perception a...

Xiang-Min Kong, Xi-Bin Jia, Da-Wei Yang et al. · 0 citations
Open access Aug 2026

Attention-Guided Vision-Language Model for Automated Radiology Report Generation

The proposed AG-VLM framework provides a scalable foundation for computer-assisted radiology reporting while retaining the need for radiologist verification before clinical use and indicates that explicit attention-guided visual reasoning combined with cross-modal semantic alignment can generate more accurate, clinical...

P. Dayaker, M. Vignesh, I. Z. et al. · 0 citations
Open access Aug 2026

Medrecord-CLIP: enhancing fundus disease diagnosis via EHR-guided vision-language pre-training

This work proposes MedRecord-CLIP, a knowledge-enhanced foundation model featuring a diagnosis-guided cross-attention mechanism to adaptively extract and fuse salient patient history with diagnostic representations that highlights the critical value of integrating personalized clinical context to enhance the generaliza...

Lei Shi, Wenbin Zhai, Lei Yu et al. · 0 citations
Preprint Aug 2026

SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

SeVeR is proposed, a selective visual exposure framework that compresses dense volumes into modality-wise prototypes and retrieves complementary multi-level evidence with change-aware gated attention during decoding, trained with a marginal-utility self-consistency objective that suppresses unhelpful retrieval.

Yao-Jun Hu, Danyang Tu, Yang Liu et al. · 0 citations
Preprint Aug 2026

MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as...

Rafi Ibn Sultan, Hui Zhu, Chengyin Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.