Skip to content

Information-Bottleneck-Guided Hybrid Neural Architecture Search for Temporal Action Detection in Untrimmed Videos

Aug 2026 · IEEE Transactions on Image Processing · Vol 35, pp. 8664-8677 · 0 citations · 75 references
Medicine Computer Science

Abstract

Temporal Action Detection (TAD) in untrimmed videos requires effective spatial feature extraction for precise action classification and temporal feature modeling for accurate boundary localization. To achieve effective spatio-temporal feature integration, several works manually design rule-based (i.e., sequential or parallel) hybrid Mamba-Transformer networks for TAD. However, few studies explore diverse integration strategies and network topologies due to the inherent limitations of manual design. Therefore, we propose NAS-TAD, the first Neural Architecture Search framework for TAD, systematically exploring this untouched problem. Specifically, we develop a spatio-temporal NAS objective function based on information-bottleneck theory to quantify task-relevant spatio-temporal features, providing interpretable guidance for the network search and optimization process. Furthermore, we reformulate Transformer self-attention as a state-space model, thereby enabling seamless switching between Mamba and Transformer blocks in a unified weight-sharing search space. Consequently, comprehensive experiments on ActivityNet, THUMOS14, HACS and FineAction demonstrate the effectiveness of the searched hybrid architectures, providing new insights into temporal and spatial feature fusion for TAD. Code is available for reproduction at https://github.com/tyhnu/nastad.git

View source

Similar papers

Spatiotemporal agent based action recognition using YOWO-former

Spatio-temporal action detection (STAD) localizes and classifies human actions in video simultaneously. YOWO is a one-stage dual-stream detector that decouples spatial (2D CNN) and temporal (3D CNN) backbones, yet all variants rely on 3D CNNs with limited temporal receptive fields. Video transformers such as VideoMAE l...

Vasu Thakaew · 0 citations
Open access 2026

Cross-Modal Temporal Alignment for Action Grounding in Videos

A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.

Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al. · 0 citations
Open access Sep 2026

Domain-Robust Temporal Action Detection for Power Grid Operation Videos via Consistency-Guided Domain Modeling

Temporal action detection (TAD) should remain reliable when videos are collected by different sources, even if the action vocabulary is unchanged. We study this domain-generalization setting with pre-extracted temporal features and no target-domain videos during training. DRTAD models source-related feature factors, mi...

Ling-Wen Meng, Guang-Hui Xi, Jian-Gang Liu et al. · 0 citations
Aug 2026

DiaVTG: multi-turn reasoning framework for video temporal grounding

DiaVTG, a novel VTG framework designed to enhance temporal localization precision, is proposed to reformulate temporal localization as a video understanding problem and demonstrates that the training-free method consistently improves performance across various Vid-LLM architectures.

Hong-Yu Huang, Junyi Yang, Sipeng Yang et al. · 0 citations
Open access Aug 2026

MSDR-Mamba: A Multi-Scale Branch-Decoupled Routing State-Space Detector for Temporal Action Localization

The results support scale- and branch-aware state-space modeling as a practical design strategy for TAL as well as scale- and branch-aware state-space modeling as a practical design strategy for scale- and branch-aware state-space modeling.

Rui-Jun Gu, Wenyan Bi, Yu Han et al. · 0 citations
Preprint Aug 2026

MASQ: Mask-Aware Spatiotemporal Quantization for Unsupervised Skeleton Action Segmentation

A novel Mask-aware Action Spatiotemporal Quantization framework that decouples the conflicting tasks of spatial feature inference and temporal smoothing, and establishes a comprehensive and substantial leading advantage in the Mean over Frames accuracy.

Xingchen Qin, Lin-Xiang Peng, You-Bao Ye et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.