Aug 2026· International Conference on Image Processing. Machine Learning and Pattern Recognition· Vol 14304, pp. 143040K - 143040K-8· 0 citations· 25 references
Engineering
TL;DR
A pose-guided multimodal textualization framework that integrates real-time skeletal pose estimation from monocular RGB video with physiological sensor streams—including respiration, trigger pressure, IMU, heart-rate variability, and eye-tracking—through a unified vision-language reasoning pipeline is proposed.
Abstract
Accurate assessment of shooting performance requires understanding of body posture, physiological state, and fine motor control. Existing approaches either focus on isolated sensor modalities or rely on opaque numerical pipelines that lack interpretability. This paper proposes a pose-guided multimodal textualization framework that integrates real-time skeletal pose estimation from monocular RGB video with physiological sensor streams—including respiration, trigger pressure, IMU, heart-rate variability, and eye-tracking—through a unified vision-language reasoning pipeline. A lightweight pose estimator extracts body-alignment features (shoulder levelness, elbow stability, center-of-mass displacement), which are jointly textualized with physiological indicators and fed to a large language model (LLM) for score prediction, factor-level reasoning, and training feedback generation. Experiments on approximately 70 athletes and 20,000 samples show that the proposed method achieves 83.7% classification accuracy and 0.80 F1-score, surpassing the strongest baseline (LSTM: 80.1%) by 3.6 percentage points while providing coach-readable explanations. Ablation analysis confirms that vision-based pose features provide complementary spatial cues to inertial measurements, yielding a 2.3 percentage point accuracy gain. Coach-rated interpretability reaches 4.6/5.0, substantially exceeding XGBoost (3.1) and LSTM (2.8).
Human motion understanding requires not only accurate action recognition but also interpretable performance evaluation capable of reflecting motion quality. This paper proposes Pose-ARPA, a unified framework that combines deep pose estimation with spatiotemporal representation learning for comprehensive action recognit...
Qi Yang, Long-Hui Wen· International Conference on...· 0 citations
Exercise recognition and movement quality assessment remain challenging in supervised exercise training, particularly under viewpoint changes and self-occlusion. Vision-based methods provide spatial posture information but are susceptible to keypoint loss, whereas inertial sensing is less affected by occlusion but prov...
Zhao-Yan Gu, Ruo-Peng Yang, Yong-Qi Shi et al.· Italian National Conference...· 0 citations
Existing VR datasets lack synchronized multimodal pairing of headset tracking and egocentric visual context, making it difficult to systematically study when and how visual cues improve head-motion prediction. While prior prediction work suggests visual information can help, these claims often rest on isolated models r...
Zi-Yu Zhong, Jens Nirme, Héctor A. Caltenco et al.· Proceedings of the 28th Inte...· 0 citations
MMAF-Net is proposed, a multi-branch deep learning architecture that integrates visual, pose-estimation and inertial sensor streams to extract complementary motion features, fused via a temporal attention module to classify 23 karate action types accurately.
Yong Gao, Zhao-Hui Liu· International Journal of e-c...· 0 citations
Vision-based badminton analysis requires the joint understanding of body posture, shuttle movement, court geometry, and stroke-level semantics. Existing studies usually address posture recognition, stroke forecasting, shuttle tracking, and coaching recommendation as separate tasks, which limits their ability to explain...