Skip to content
Conference

Efficient Frame Retrieval for Traffic Law Question Answering over Dashcam Videos

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 730-735 · 0 citations · 24 references

Abstract

Traffic Law Question Answering over dashcam videos requires both accurate visual evidence selection and reliable multimodal reasoning. This task is especially challenging in Vietnamese traffic scenes, where road environments are dense, diverse, and highly dynamic. In this paper, we present a resource-efficient two-stage framework for traffic-law question answering over dashcam videos, developed for the "RoadBuddy: Understanding the Road through Dashcam AI" track of ZaloAI Challenge 2025. First, a CLIP ViT–based frame retriever reduces visual redundancy by selecting four query-relevant frames from each video. Second, the selected frames are processed by a Qwen3-VL 8B model adapted with LoRA for task-specific multiple-choice answer prediction. Despite its lightweight design, our method achieves competitive results, with scores of 0.70617 on the public leaderboard and 0.702 on the private leaderboard, ranking top 3 among 241 participating teams. These results show that efficient frame retrieval, when combined with parameter-efficient task adaptation, offers a practical solution for traffic-law question answering over real-world dashcam videos.

View source

Similar papers

Preprint Aug 2026

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

UniTraffic-Agent is introduced, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning.

Peng Li, Qianqian Xu, Shilong Bao et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Query-aligned video frame selection for long video understanding

Multimodal large language models (MLLMs) process multimodal inputs by converting text, images, and videos into token sequences that are subsequently processed by a backbone language model. While MLLMs have achieved excellent performance in understanding the content of individual images, video understanding remains sign...

Md. Safayet Islam, Dilip Sarkar, Liang-Yu Liang · 0 citations
Preprint Sep 2026

Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA

A decoupled semantic understanding framework that first resolves predefined traffic questions into structured semantic facts and subsequently uses these facts to guide caption generation, leading to more accurate and reliable traffic understanding, leading to higher-quality lan guage generation.

Bui Hoai Thuong Nguyen, Thanh-Nhan Vo, T. Nguyen et al. · 1 citation
Preprint Aug 2026

Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering

Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input....

Hsiang-Wei Huang, Fu-Chen Chen, Li-Wu Tsao et al. · 0 citations
Conference Open access Sep 2026

Dynamic Multi-Path Retrieval for Knowledge-based Visual Question Answering

Dynamic Multi-Path Retrieval for KB-VQA (DMRAG) is proposed, which re-trieves candidates through multiple retrieval paths that capture complementary visual and semantic cues and performs Question-Adaptive Gated Fusion to balance contributions from different modalities according to the query’s information need.

Zeyu Song, Yimin Deng, Yu-Xin Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.