Vision-Based Skeleton–Shuttle–Court Fusion for Badminton Posture Recognition and Coaching-Oriented Performance Recommendation
Abstract
Vision-based badminton analysis requires the joint understanding of body posture, shuttle movement, court geometry, and stroke-level semantics. Existing studies usually address posture recognition, stroke forecasting, shuttle tracking, and coaching recommendation as separate tasks, which limits their ability to explain whether an action is technically executable, spatially valid, and physically economical. This paper proposes SST-CoachNet, a skeleton–shuttle–court fusion framework for badminton posture recognition and coaching-oriented performance recommendation. The framework represents badminton performance through three complementary information streams: visual skeletal posture, shuttle-motion cues, and structured court–stroke semantics. A constraint-aware hierarchical hybrid-action generator is designed to infer coaching intention, stroke category, shuttle landing point, and recovery position. Court-boundary, posture-execution, and movement-load constraints are incorporated to reduce unrealistic recommendations. A multi-criteria coaching reward model is further developed to evaluate actions from outcome contribution, initiative creation, opponent displacement, posture feasibility, spatial risk, and recovery cost. Because large-scale fully aligned video–stroke multimodal badminton data remain limited, experiments are conducted using a main two-track protocol supplemented by a manually aligned end-to-end pilot subset. The visual branch is evaluated on badminton video clips and achieves 88.9% stroke accuracy and 87.4% posture-quality F1-score. The structured recommendation branch obtains a coaching score of 0.921 with an invalid recommendation rate of 2.0%, outperforming Seq-Transformer, RallyNet, SPAIT, Decision Transformer, CQL-based recommendation, and a Flat Hybrid Generator. Compared with the strongest structured baseline, SPAIT, the coaching score increases from 0.834 to 0.921 and the invalid recommendation rate decreases from 3.5% to 2.0%. In the end-to-end aligned subset, the full detected-vision pipeline achieves a coaching score of 0.893, showing that visual posture and shuttle-tracking errors remain controllable when propagated to downstream recommendation. These results suggest that integrating visual biomechanics, shuttle-motion cues, and structured court semantics can support interpretable and practically valid AI-assisted badminton coaching.