Skip to content
Preprint

Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA

Sep 2026 · 1 citation · 25 references
Computer Science

TL;DR

A decoupled semantic understanding framework that first resolves predefined traffic questions into structured semantic facts and subsequently uses these facts to guide caption generation, leading to more accurate and reliable traffic understanding, leading to higher-quality lan guage generation.

Abstract

Track 2 of the AI City Challenge 2026 requires both visual question answering (VQA) and traffic event description generation under a challenging synthetic-to real domain shift. Existing vision-language approaches often entangle semantic understanding with language generation, making them susceptible to hallucination and inconsistent reasoning across event phases. In this work, we propose a decoupled semantic understanding framework that first resolves predefined traffic questions into structured semantic facts and subsequently uses these facts to guide caption generation. A frozen V-JEPA encoder extracts predictive scene representations, while a lightweight Llama-based predictor produces answers for VQA queries. To improve reliability, we introduce a training-free structured refinement mechanism that exploits statistical priors, inter-question relationships, and temporal event consistency to correct prediction errors. The refined semantic facts are then provided to Qwen3-VL-8B to generate pedestrian and vehicle descriptions for each traffic event. Experimental results on the official 2026 AI City Challenge Track 2 benchmark show that the proposed method achieves 87.09% VQA accuracy and an overall S2 score of 60.0853, ranking first among all participating teams. These results demonstrate that predictive world representations combined with structured semantic refinement enable more accurate and reliable traffic understanding, leading to higher-quality lan guage generation.

View source

Similar papers

Preprint Aug 2026

VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction

The results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary, and this model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies.

Yu-Chen Zhang, Yuan Gao, Sebastian Schmidt et al. · 0 citations
Preprint Aug 2026

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

UniTraffic-Agent is introduced, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning.

Peng Li, Qianqian Xu, Shilong Bao et al. · 1 citation
Preprint Sep 2026

CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning

Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle's field of view, yet conventional cooperative-driving systems typically transform them into geomet...

Kang Yang, Shuai-Yu Liu, Han-Sen Li et al. · 0 citations
#machine learning Preprint Sep 2026

Grounding Vision-Language Models in Driving Semantics: A Multi-Dataset Predicate Framework for Explainable Reasoning

Vision-language models are increasingly used for driving-scene understanding, yet the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation. This paper introduces a deterministic multi-dataset predicate framework that derives driving-scene semantics from me...

Mohamed Chouai, Fazli Faruk Okumus, Stefan Kugele · 0 citations
Review Sep 2026

Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models

Automated football news generation from raw videos requires bridging spatiotemporal perception with factual text composition. This study develops an end-to-end, event-based framework that converts match videos into fact-grounded reports. The framework uses an Inflated Three-Dimensional ConvNet (I3D) backbone with multi...

Yi-Feng Wang, Yi-Hang Huang · 0 citations
Preprint Aug 2026

Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum

This work builds a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training and proposes CurrSTVG, a curriculum reinforcement learning...

Xing-Jian Wang, Shijian Wang, Yi-Bo Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.