Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.
Abstract
Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions among these abilities under-explored. We introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten tasks in text, image, and video. It is built from the 2025-2026 NBA season and includes official playby-play, rosters and profiles for 530 active players, and 2,501 possession-level broadcast clips. We further propose BasketballSkills, an agent that composes eight basketball-specific perception and retrieval tools under four reusable skills that specify tool order, evidence bindings, and stopping conditions. Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.
Visual understanding in sports has emerged as a hot topic in computer vision in recent years. Most existing basketball video datasets adopt single action or activity as sample, which can neither preserve the temporal continuity of game events nor support complex tasks such as action anticipation. To address this issue,...
Yunhao Zhao, Haoying Sun, Jiarui Li et al.· 0 citations
Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Amateur team sport is a useful, largely untested place to check that assumption: over eight million students played a schoo...
Sai Varun Kodathala, Prashanth Pollishetty, J. Cargill· 0 citations
Introduction & Purpose
Game models – integrated collections of principles that determine a desired team identity and player behaviours – are paramount in modern elite football coaching. In practice, quantitative methods typically operate on untested ontologies derived from pre-existing game models (Corsie et al., 2024;...
Jonas Bischofberger, Run-Qing Ma, Arnold Baca· Current Issues in Sport Scie...· 0 citations
Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perceive individual strokes without modeling their tactical dependencies or generate high-level analyses without grounding them in the underlyin...
Yifan Mei, Qin-Ling Shi, Chang-Li Wu et al.· 0 citations
Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for eval...
Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini et al.· 0 citations
Automated football news generation from raw videos requires bridging spatiotemporal perception with factual text composition. This study develops an end-to-end, event-based framework that converts match videos into fact-grounded reports. The framework uses an Inflated Three-Dimensional ConvNet (I3D) backbone with multi...
Yi-Feng Wang, Yi-Hang Huang· International journal of pat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.