Skip to content
Preprint

Towards Comprehensive Basketball Understanding

Aug 2026 · 1 citation · 37 references
Computer Science

TL;DR

Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.

Abstract

Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions among these abilities under-explored. We introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten tasks in text, image, and video. It is built from the 2025-2026 NBA season and includes official playby-play, rosters and profiles for 530 active players, and 2,501 possession-level broadcast clips. We further propose BasketballSkills, an agent that composes eight basketball-specific perception and retrieval tools under four reusable skills that specify tool order, evidence bindings, and stopping conditions. Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.

View source

Similar papers

Preprint Aug 2026

PL-NBA: A Possession-level Universal Basketball Video Dataset Supporting Multiple Visual Understanding Tasks

Visual understanding in sports has emerged as a hot topic in computer vision in recent years. Most existing basketball video datasets adopt single action or activity as sample, which can neither preserve the temporal continuity of game events nor support complex tasks such as action anticipation. To address this issue,...

Yunhao Zhao, Haoying Sun, Jiarui Li et al. · 0 citations
#computer vision Preprint Sep 2026

Prompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings

Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Amateur team sport is a useful, largely untested place to check that assumption: over eight million students played a schoo...

Sai Varun Kodathala, Prashanth Pollishetty, J. Cargill · 0 citations
Open access Sep 2026

Evaluating the “Players First” Game Model Using Large-Scale Tracking Data

Introduction & Purpose Game models – integrated collections of principles that determine a desired team identity and player behaviours – are paramount in modern elite football coaching. In practice, quantitative methods typically operate on untested ontologies derived from pre-existing game models (Corsie et al., 2024;...

Jonas Bischofberger, Run-Qing Ma, Arnold Baca · 0 citations
Preprint Aug 2026

TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos

Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perceive individual strokes without modeling their tactical dependencies or generate high-level analyses without grounding them in the underlyin...

Yifan Mei, Qin-Ling Shi, Chang-Li Wu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence

Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for eval...

Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini et al. · 0 citations
Review Sep 2026

Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models

Automated football news generation from raw videos requires bridging spatiotemporal perception with factual text composition. This study develops an end-to-end, event-based framework that converts match videos into fact-grounded reports. The framework uses an Inflated Three-Dimensional ConvNet (I3D) backbone with multi...

Yi-Feng Wang, Yi-Hang Huang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.