Skip to content
Conference

Controllable medical report generation via multiscale dynamic focusing and instruction fine-tuning

Sep 2026 · International Conference on Artificial Intelligence, Machine, Vision and Control · Vol 14345, pp. 143451R - 143451R-12 · 0 citations · 41 references
Engineering

Abstract

Radiologists typically adopt a coarse-to-fine, dynamic focusing cognitive strategy when interpreting medical images, focusing on potential abnormal regions while integrating semantic context to compose reports. However, most existing medical report generation methods rely on fixed-resolution image encoding, which struggle to accommodate the radiologist's spatial interaction intent, resulting in insufficient fine-grained lesion descriptions. To address this, we propose a controllable medical report generation framework centered on physician interaction. The framework introduces the mechanism of "Select bbox-to-Explain", utilizing physician-specified spatial regions as instruction signals. The model constructs high-resolution local views and global contextual representations through a bounding-box-guided multi-scale dynamic focusing mechanism. During the report generation phase, regional visual features are parsed into structured medical semantics, which are then modeled alongside spatial instructions and natural language task prompts as a unified Multimodal Prompt. This architecture drives a Large Language Model (LLM) to generate controllable and clinically consistent medical reports. Experimental results demonstrate that the proposed model achieves a ROUGE-L score of 0.402 on region-level reports—an approximate 40% improvement over the baseline while maintaining the contextual integrity of the full report. This highlights the model's efficacy in enhancing critical region representation and semantic consistency, thereby improving clinical trust in computer-aided diagnosis.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.