SCAN-R: bridging precision imaging and natural language for automated stroke diagnosis
Abstract
Stroke is the second leading cause of death globally, where each minute of treatment delay results in the loss of 1.9 million brain cells. Traditional CT-based diagnosis relies on manual interpretation with inherent variability and time constraints, while existing AI approaches typically address only isolated tasks such as detection or segmentation without integrated clinical reporting. We present Stroke CT Analysis and Natural Language Reporting (SCAN-R), a unified end-to-end framework that integrates multiclass stroke detection, Transformerenhanced U-Net segmentation with task-specific pre-trained backbones, and Retrieval-Augmented Generation for evidence-based clinical report generation. Evaluation on 6,653 CT scans demonstrates 95.81% detection accuracy and a Dice coefficient of 0.81 for bleeding lesion segmentation, and 10% improvement in clinical decision-making quality on the MedMCQA benchmark. The framework successfully transforms raw CT images into structured clinical reports with quantitative metadata and evidence-based recommendations, demonstrating potential to accelerate time-critical stroke diagnosis in emergency settings.