Spatiotemporal agent based action recognition using YOWO-former
Abstract
Spatio-temporal action detection (STAD) localizes and classifies human actions in video simultaneously. YOWO is a one-stage dual-stream detector that decouples spatial (2D CNN) and temporal (3D CNN) backbones, yet all variants rely on 3D CNNs with limited temporal receptive fields. Video transformers such as VideoMAE learn stronger temporal representations but produce token-sequence outputs incompatible with CNN-based detectors. This paper presents YOWOFormer, which replaces the 3D CNN with VideoMAE and adopts YOLOv11 as the 2D backbone, introducing two bridging modules: (1) a Spatial Query Adapter (SQA) that converts transformer tokens into spatial feature maps via learnable queries and cross-attention, and (2) a Bidirectional Spatio-Temporal Fusion (BSTF) module that enables mutual reinforcement between spatial and temporal features through bidirectional cross-attention. We systematically compare four adapter designs, three fusion strategies, and the effects of scaling backbone capacity and clip length. On UCF101-24, YOWOFormer-L achieves 93.35% frame-level mAP, surpassing YOWOv3-L (88.64%), while YOWOFormer-B attains 91.47% at 26.81 FPS. On AVA v2.2, YOWOFormer-B and -L improve by 2.34 and 5.39 points over prior 3D-CNN-based dual-stream methods. Ablation studies confirm that scaling video transformers yields greater accuracy gains than scaling 3D CNNs.