Skip to content
Preprint

Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum

Aug 2026 · 0 citations · 35 references
Computer Science

TL;DR

This work builds a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training and proposes CurrSTVG, a curriculum reinforcement learning framework that delivers consistent gains.

Abstract

Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely overlook the compositional queries prevalent in real-world scenarios, where a target must be disambiguated by jointly reasoning about its attributes and relations to other entities. To bridge this gap, we propose Compositional Spatio-Temporal Video Grounding (CompSTVG), a task that requires models to process complex textual queries where every intertwined attribute and relational cue is essential for disambiguation. To facilitate this task at scale, we build a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training. Built on this engine, we introduce STVG-CompBench, a benchmark stratified by explicit difficulty levels that jointly capture temporal complexity and spatial interference. Evaluating 11 representative STVG models on STVG-CompBench reveals that current models perform poorly on compositional queries, exhibiting a sharp performance drop that is typically obscured by overall dataset-level averages. We further construct synthetic training data and propose CurrSTVG, a curriculum reinforcement learning framework that delivers consistent gains, with the largest improvements observed on the most challenging compositional queries.

View source

Similar papers

Preprint Aug 2026

ID-VTG: Image-Disambiguated Video Temporal Grounding

The Visually-Guided Disambiguation Aggregation Aggregation (VGD-Agg) framework is proposed, a framework based on a dual-branch fast-slow architecture that enhances discriminability via two learnable tokens and achieves state-of-the-art results on the proposed benchmarks.

Minghang Zheng, Jing Wei, Hong-Yi Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding

Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query...

Nan-Xing Hu, Xiao-Yue Duan, Qi-Wei Yan et al. · 0 citations
Preprint Aug 2026

Modeling Scientific Experiment Scenes: Dataset and Model

The Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG that enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues, is proposed.

Ming-Hao Zou, Qingtian Zeng, Shangkun Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding

Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typically rely on computationally expensive architectures, complex t...

Alberto Presta, Michal Byra, Grzegorz Stefański et al. · 0 citations
#computer vision Preprint Aug 2026

STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models

STRAND is introduced, a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness and an object-centric framework that explicitly constructs and reasons over structured obj...

T. Nguyen, Tri Cao, Khoi M. Le et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBe...

Hyungjin Chung, Byeongjun Park, Joonseok Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.