A decoupled semantic understanding framework that first resolves predefined traffic questions into structured semantic facts and subsequently uses these facts to guide caption generation, leading to more accurate and reliable traffic understanding, leading to higher-quality lan guage generation.
Abstract
Track 2 of the AI City Challenge 2026 requires both visual question answering (VQA) and traffic event description generation under a challenging synthetic-to real domain shift. Existing vision-language approaches often entangle semantic understanding with language generation, making them susceptible to hallucination and inconsistent reasoning across event phases. In this work, we propose a decoupled semantic understanding framework that first resolves predefined traffic questions into structured semantic facts and subsequently uses these facts to guide caption generation. A frozen V-JEPA encoder extracts predictive scene representations, while a lightweight Llama-based predictor produces answers for VQA queries. To improve reliability, we introduce a training-free structured refinement mechanism that exploits statistical priors, inter-question relationships, and temporal event consistency to correct prediction errors. The refined semantic facts are then provided to Qwen3-VL-8B to generate pedestrian and vehicle descriptions for each traffic event. Experimental results on the official 2026 AI City Challenge Track 2 benchmark show that the proposed method achieves 87.09% VQA accuracy and an overall S2 score of 60.0853, ranking first among all participating teams. These results demonstrate that predictive world representations combined with structured semantic refinement enable more accurate and reliable traffic understanding, leading to higher-quality lan guage generation.
The results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary, and this model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies.
Yu-Chen Zhang, Yuan Gao, Sebastian Schmidt et al.· 0 citations
UniTraffic-Agent is introduced, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning.
Peng Li, Qianqian Xu, Shilong Bao et al.· 1 citation
Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle's field of view, yet conventional cooperative-driving systems typically transform them into geomet...
Kang Yang, Shuai-Yu Liu, Han-Sen Li et al.· 0 citations
Vision-language models are increasingly used for driving-scene understanding, yet the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation. This paper introduces a deterministic multi-dataset predicate framework that derives driving-scene semantics from me...
Mohamed Chouai, Fazli Faruk Okumus, Stefan Kugele· 0 citations
Automated football news generation from raw videos requires bridging spatiotemporal perception with factual text composition. This study develops an end-to-end, event-based framework that converts match videos into fact-grounded reports. The framework uses an Inflated Three-Dimensional ConvNet (I3D) backbone with multi...
Yi-Feng Wang, Yi-Hang Huang· International journal of pat...· 0 citations
This work builds a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training and proposes CurrSTVG, a curriculum reinforcement learning...
Xing-Jian Wang, Shijian Wang, Yi-Bo Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.