Multimodal Radiology Assistant with Graph- Enhanced Reasoning and Uncertainty-Guided Report Generation
Abstract
Automatic radiology report generation has become an active research area due to its potential to reduce radiologist workload and standardize reporting quality. However, state-of- the-art systems still suffer from hallucinated findings, limited clinical reasoning, and a lack of calibrated uncertainty estimates, all of which undermine trust in realworld deployments. This paper presents a clinically-aware multimodal radiology assis- tant that combines graph-enhanced cross-modal reasoning with uncertainty-guided report generation. The assistant integrates a convolutional or vision transformer backbone with a structured clinical knowledge graph, coupled through a graph-enhanced cross-attention module to align visual features with anatomical and pathological entities. A clinical consistency loss penalizes mismatches between image-based abnormality predictions and findings expressed in the generated text, directly targeting hallucinations. Uncertainty is quantified using Monte Carlo dropout and temperature scaling, enabling confidence-aware outputs and uncertainty-guided triage. Experiments on two public chest X- ray benchmarks and cross-dataset evaluation demonstrate im- proved factual consistency, reduced hallucination rate, and better calibration, including in rare-disease subsets, compared with transformer-only baselines. These results suggest that structured reasoning and uncertainty modeling are key ingredients for trustworthy radiology report generation.