Aug 2026· International Conference on Multimedia Analysis and Pattern Recognition· pp. 730-735· 0 citations· 24 references
Abstract
Traffic Law Question Answering over dashcam videos requires both accurate visual evidence selection and reliable multimodal reasoning. This task is especially challenging in Vietnamese traffic scenes, where road environments are dense, diverse, and highly dynamic. In this paper, we present a resource-efficient two-stage framework for traffic-law question answering over dashcam videos, developed for the "RoadBuddy: Understanding the Road through Dashcam AI" track of ZaloAI Challenge 2025. First, a CLIP ViT–based frame retriever reduces visual redundancy by selecting four query-relevant frames from each video. Second, the selected frames are processed by a Qwen3-VL 8B model adapted with LoRA for task-specific multiple-choice answer prediction. Despite its lightweight design, our method achieves competitive results, with scores of 0.70617 on the public leaderboard and 0.702 on the private leaderboard, ranking top 3 among 241 participating teams. These results show that efficient frame retrieval, when combined with parameter-efficient task adaptation, offers a practical solution for traffic-law question answering over real-world dashcam videos.
UniTraffic-Agent is introduced, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning.
Peng Li, Qianqian Xu, Shilong Bao et al.· 1 citation
Multimodal large language models (MLLMs) process multimodal inputs by converting text, images, and videos into token sequences that are subsequently processed by a backbone language model. While MLLMs have achieved excellent performance in understanding the content of individual images, video understanding remains sign...
TAR and TAR-Bench serve as the official training and in-domain evaluation resources for AI City Challenge 2026 Track 3 and highlight persistent limitations in temporal precision and causal attribution.
Han Zhang, Yi-Lin Zhao, Zaid Pervaiz Bhat et al.· 1 citation
A decoupled semantic understanding framework that first resolves predefined traffic questions into structured semantic facts and subsequently uses these facts to guide caption generation, leading to more accurate and reliable traffic understanding, leading to higher-quality lan guage generation.
Bui Hoai Thuong Nguyen, Thanh-Nhan Vo, T. Nguyen et al.· 1 citation
Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input....
Hsiang-Wei Huang, Fu-Chen Chen, Li-Wu Tsao et al.· 0 citations
Dynamic Multi-Path Retrieval for KB-VQA (DMRAG) is proposed, which re-trieves candidates through multiple retrieval paths that capture complementary visual and semantic cues and performs Question-Adaptive Gated Fusion to balance contributions from different modalities according to the query’s information need.
Zeyu Song, Yimin Deng, Yu-Xin Zhang et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.