Vision-augmented multimodal analysis for competitive rifle shooting via pose-guided textualization and LLM reasoning
A pose-guided multimodal textualization framework that integrates real-time skeletal pose estimation from monocular RGB video with physiological sensor streams—including respiration, trigger pressure, IMU, heart-rate variability, and eye-tracking—through a unified vision-language reasoning pipeline is proposed.