Aug 2026· IEEE Transactions on Image Processing· Vol 35, pp. 8664-8677· 0 citations· 75 references
MedicineComputer Science
Abstract
Temporal Action Detection (TAD) in untrimmed videos requires effective spatial feature extraction for precise action classification and temporal feature modeling for accurate boundary localization. To achieve effective spatio-temporal feature integration, several works manually design rule-based (i.e., sequential or parallel) hybrid Mamba-Transformer networks for TAD. However, few studies explore diverse integration strategies and network topologies due to the inherent limitations of manual design. Therefore, we propose NAS-TAD, the first Neural Architecture Search framework for TAD, systematically exploring this untouched problem. Specifically, we develop a spatio-temporal NAS objective function based on information-bottleneck theory to quantify task-relevant spatio-temporal features, providing interpretable guidance for the network search and optimization process. Furthermore, we reformulate Transformer self-attention as a state-space model, thereby enabling seamless switching between Mamba and Transformer blocks in a unified weight-sharing search space. Consequently, comprehensive experiments on ActivityNet, THUMOS14, HACS and FineAction demonstrate the effectiveness of the searched hybrid architectures, providing new insights into temporal and spatial feature fusion for TAD. Code is available for reproduction at https://github.com/tyhnu/nastad.git
Spatio-temporal action detection (STAD) localizes and classifies human actions in video simultaneously. YOWO is a one-stage dual-stream detector that decouples spatial (2D CNN) and temporal (3D CNN) backbones, yet all variants rely on 3D CNNs with limited temporal receptive fields. Video transformers such as VideoMAE l...
A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.
Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al.· IEEE Access· 0 citations
Temporal action detection (TAD) should remain reliable when videos are collected by different sources, even if the action vocabulary is unchanged. We study this domain-generalization setting with pre-extracted temporal features and no target-domain videos during training. DRTAD models source-related feature factors, mi...
Ling-Wen Meng, Guang-Hui Xi, Jian-Gang Liu et al.· Electronics· 0 citations
DiaVTG, a novel VTG framework designed to enhance temporal localization precision, is proposed to reformulate temporal localization as a video understanding problem and demonstrates that the training-free method consistently improves performance across various Vid-LLM architectures.
Hong-Yu Huang, Junyi Yang, Sipeng Yang et al.· The Visual Computer· 0 citations
The results support scale- and branch-aware state-space modeling as a practical design strategy for TAL as well as scale- and branch-aware state-space modeling as a practical design strategy for scale- and branch-aware state-space modeling.
Rui-Jun Gu, Wenyan Bi, Yu Han et al.· Electronics· 0 citations
A novel Mask-aware Action Spatiotemporal Quantization framework that decouples the conflicting tasks of spatial feature inference and temporal smoothing, and establishes a comprehensive and substantial leading advantage in the Mean over Frames accuracy.
Xingchen Qin, Lin-Xiang Peng, You-Bao Ye et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.