Skip to content
Conference

STGraphVQA: spatial-temporal graph reasoning with hierarchical cognition for interpretable driving scene understanding

Jul 2026 · International Conference on Image Processing and Intelligent Control · Vol 14262, pp. 142621I - 142621I-7 · 0 citations · 9 references
Engineering

Abstract

Visual question answering (VQA) in autonomous driving scenarios demands strong spatiotemporal reasoning capabilities, yet existing vision-language models lack explicit modeling of dynamic relationships in complex traffic scenes. We propose STGraphVQA, a framework that represents driving scenes as dynamic spatiotemporal graphs, where nodes denote traffic participants, edges encode spatial and semantic relationships, and the temporal dimension captures their evolution. A hierarchical reasoning architecture progressively processes information through perception, relation, and decision layers, simulating the human driving cognitive process. A logit-level constrained decoding mechanism further ensures that generated answers comply with traffic rules and physical feasibility. Experiments on DriveLM and STRIDE-QA demonstrate that STGraphVQA significantly outperforms state-of-the-art baselines, achieving a Top-1 accuracy of 76.8% and a reasoning chain completeness of 82.3%, providing a promising direction toward interpretable autonomous driving systems.

View source