Skip to content
Conference

Cross-modal bidirectional gating for pose-enhanced video action detection

Jul 2026 · International Conference on Machine Vision, Automatic Identification and Detection · Vol 14261, pp. 142610P - 142610P-6 · 0 citations · 14 references
Engineering

TL;DR

BPMF is proposed, a mutual gating module that enables pose features to guide RGB attention toward action-critical regions while RGB features reweight pose responses to suppress unreliable structural cues and achieves superior results on JHMDB21.

Abstract

Video action detection requires simultaneous actor localization and action recognition across temporal sequences. Although recent RGB-based methods have achieved strong performance, they often struggle with appearance ambiguity, background clutter, and occlusion. Human pose provides complementary structural cues that are semantically meaningful and less sensitive to irrelevant background information. However, existing RGB-pose fusion strategies are typically oneway, using pose only to guide visual features while ignoring the fact that RGB appearance can also help assess the reliability of pose representations. In this paper, we propose Bidirectional Pose-RGB Modulation Fusion (BPMF), a mutual gating module that enables pose features to guide RGB attention toward action-critical regions while RGB features reweight pose responses to suppress unreliable structural cues. We integrate BPMF into a Mamba-based video action detector with a lightweight pose branch for 2D skeleton modeling. Experiments on JHMDB21 show that our method achieves 76.96% frame-mAP, outperforming both the RGB-only baseline by 4.00 percentage points and the unidirectional fusion variant by 2.84 percentage points. Ablation studies further confirm the effectiveness of the proposed bidirectional design.

View source

Similar papers

Open access 2026

Cross-Modal Temporal Alignment for Action Grounding in Videos

A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.

Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al. · 0 citations
Conference Sep 2026

Deep pose estimation-based action recognition and performance analysis for intelligent motion understanding

Human motion understanding requires not only accurate action recognition but also interpretable performance evaluation capable of reflecting motion quality. This paper proposes Pose-ARPA, a unified framework that combines deep pose estimation with spatiotemporal representation learning for comprehensive action recognit...

Qi Yang, Long-Hui Wen · 0 citations
Conference Aug 2026

Vision-augmented multimodal analysis for competitive rifle shooting via pose-guided textualization and LLM reasoning

A pose-guided multimodal textualization framework that integrates real-time skeletal pose estimation from monocular RGB video with physiological sensor streams—including respiration, trigger pressure, IMU, heart-rate variability, and eye-tracking—through a unified vision-language reasoning pipeline is proposed.

Wen Tao, Yan Li, Lijun Fu · 0 citations
2026

HiPATrack: Hierarchical Dependency and Position-Aware for TIR Object Tracking

Most current thermal infrared (TIR) object tracking methods are adapted from the RGB domain, and their feature fusion strategies typically emphasize high-level semantic representations. As a result, the already sparse low- and mid-level discriminative cues (e.g., texture and boundaries) in TIR images are further attenu...

Yue-Hao Li, Wei-Sheng Li, Shun-Ping Chen et al. · 0 citations
Preprint Aug 2026

Foundational feature fusion for conditional flow matching in 6D pose estimation

This work presents FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoders supervised on object-scene overlap and reducing supervision requirements and memory overhead.

Amir Hamza, Davide Boscaini, Fabio Poiesi · 0 citations
Preprint Aug 2026

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

This work introduces a weakly-supervised vision-language pretraining mechanism that transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without tra...

Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.