Jul 2026· International Conference on Machine Vision, Automatic Identification and Detection· Vol 14261, pp. 142610P - 142610P-6· 0 citations· 14 references
Engineering
TL;DR
BPMF is proposed, a mutual gating module that enables pose features to guide RGB attention toward action-critical regions while RGB features reweight pose responses to suppress unreliable structural cues and achieves superior results on JHMDB21.
Abstract
Video action detection requires simultaneous actor localization and action recognition across temporal sequences. Although recent RGB-based methods have achieved strong performance, they often struggle with appearance ambiguity, background clutter, and occlusion. Human pose provides complementary structural cues that are semantically meaningful and less sensitive to irrelevant background information. However, existing RGB-pose fusion strategies are typically oneway, using pose only to guide visual features while ignoring the fact that RGB appearance can also help assess the reliability of pose representations. In this paper, we propose Bidirectional Pose-RGB Modulation Fusion (BPMF), a mutual gating module that enables pose features to guide RGB attention toward action-critical regions while RGB features reweight pose responses to suppress unreliable structural cues. We integrate BPMF into a Mamba-based video action detector with a lightweight pose branch for 2D skeleton modeling. Experiments on JHMDB21 show that our method achieves 76.96% frame-mAP, outperforming both the RGB-only baseline by 4.00 percentage points and the unidirectional fusion variant by 2.84 percentage points. Ablation studies further confirm the effectiveness of the proposed bidirectional design.
A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.
Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al.· IEEE Access· 0 citations
Human motion understanding requires not only accurate action recognition but also interpretable performance evaluation capable of reflecting motion quality. This paper proposes Pose-ARPA, a unified framework that combines deep pose estimation with spatiotemporal representation learning for comprehensive action recognit...
Qi Yang, Long-Hui Wen· International Conference on...· 0 citations
A pose-guided multimodal textualization framework that integrates real-time skeletal pose estimation from monocular RGB video with physiological sensor streams—including respiration, trigger pressure, IMU, heart-rate variability, and eye-tracking—through a unified vision-language reasoning pipeline is proposed.
Wen Tao, Yan Li, Lijun Fu· International Conference on...· 0 citations
Most current thermal infrared (TIR) object tracking methods are adapted from the RGB domain, and their feature fusion strategies typically emphasize high-level semantic representations. As a result, the already sparse low- and mid-level discriminative cues (e.g., texture and boundaries) in TIR images are further attenu...
Yue-Hao Li, Wei-Sheng Li, Shun-Ping Chen et al.· IEEE Transactions on Automat...· 0 citations
This work presents FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoders supervised on object-scene overlap and reducing supervision requirements and memory overhead.
Amir Hamza, Davide Boscaini, Fabio Poiesi· 0 citations
This work introduces a weakly-supervised vision-language pretraining mechanism that transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without tra...
Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.