Spatiotemporal agent based action recognition using YOWO-former
Spatio-temporal action detection (STAD) localizes and classifies human actions in video simultaneously. YOWO is a one-stage dual-stream detector that decouples spatial (2D CNN) and temporal (3D CNN) backbones, yet all variants rely on 3D CNNs with limited temporal receptive fields. Video transformers such as VideoMAE l...