The results suggest that global temporal context and explicit player-position cues are useful for generating more detailed soccer commentaries, while also revealing remaining trade-offs in precise timestamp localization.
Abstract
Sports video captioning plays an important role in understanding and describing dynamic visual content, and it is particularly challenging for long soccer broadcasts. Existing dense captioning methods often rely on local clips and may miss match-level temporal context or player-position cues, which can lead to captions with insufficient spatial and event-sequence detail. To address these challenges, we propose Spatio-Contextual Perceiver Resampler (SCPR), a framework for single-anchored dense soccer video captioning. SCPR combines three components: (1) global-local feature integration for incorporating long-range video context around each caption anchor; (2) spatial position encoding based on the absolute 2D locations of players; and (3) cross-modal feature fusion that aligns global, local, and spatial features with a trainable language decoder. Evaluated under the SoccerNet-Caption protocol, SCPR improves several captioning metrics, especially CIDEr, and increases recall and F1 in the spotting stage, while some strict localization and ROUGE-style metrics remain stronger for competing baselines. These results suggest that global temporal context and explicit player-position cues are useful for generating more detailed soccer commentaries, while also revealing remaining trade-offs in precise timestamp localization.
Recent advances in computer vision have made broadcast sports videos increasingly useful for event analysis, performance assessment, and player-safety applications. In soccer, however, header spotting remains a challenging problem due to the subtle and short-lived nature of header events. This paper focuses on soccer h...
Ahmed Endris Hasen, Muhammad Shahzad Khan, Nikolaos Passalis et al.· 0 citations
Game State Reconstruction (GSR) aims to recover the spatial positions and semantic identities of all players from broadcast soccer videos, requiring robust localization, tracking, and identity association under unconstrained camera motion and severe visual ambiguity. In this work, we present Broadcast2Pitch++, an exten...
Yewon Hwang, Yin May Oo, Muhammad Amrulloh Robbani et al.· Machine Vision and Applicati...· 0 citations
Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main limitations: (1) the scarcity of high-quality public datasets with fine-grained audio-visual joint captions; (2) reinforcement-learning methods that rely on coarse rew...
Ming-Yang Wu, Kai-Tuo Feng, Bohao Li et al.· 0 citations
DiaVTG, a novel VTG framework designed to enhance temporal localization precision, is proposed to reformulate temporal localization as a video understanding problem and demonstrates that the training-free method consistently improves performance across various Vid-LLM architectures.
Hong-Yu Huang, Junyi Yang, Sipeng Yang et al.· The Visual Computer· 0 citations
Sports video analysis is crucial for athletic analytics and broadcasting enhancement. Dense sports video reasoning, however, demands a fine-grained understanding of numerous small-scale, highly interactive, and visually homogeneous entities (e.g., players sharing identical uniforms, the ball) across long temporal conte...
Yizhi Li, Jiawei Jiang, Guan-Hong Wang et al.· 0 citations
This work introduces StrAD, a benchmark for long-form AD generation on full-length videos spanning diverse genres such as movies, documentaries, short films, performances, and video games, and reformulates AD generation as streaming dense video captioning, generating ADs on the fly without ground-truth timestamps.
Julian Spravil, Sebastian Houben, Sven Behnke· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.