Design and simulation research of an intelligent stage interaction system based on real time motion capture
Abstract
Intelligent stage interaction enhances immersion by translating performers' movements into multimedia directives. Existing systems relying on wearable sensors or depth cameras face limitations like high costs, constrained movement, and sensitivity to complex lighting. Traditional vision methods also struggle with multi-person occlusion and translating low-level coordinates into high-level artistic semantics. To overcome these bottlenecks, this paper proposes a lightweight, end-to-end intelligent stage interaction system based on the YOLOv11n-Pose model. We construct a three-layer Perception-Interaction-Execution architecture with a hierarchical physical-semantic feature quantization model to dynamically map skeletal signals to artistic semantics. Additionally, the system employs a rule engine with a dual-threshold hysteresis comparator and non-linear mapping, integrated with the OSC protocol for low-latency decoupled deployment. In real-world rehearsals with natural lighting and five-person collaborative dance, the system achieves 43.1 ms latency and a 32.6 Interaction Quality Index (IQI) on edge devices, outperforming YOLOv8n. Ablation studies show the hysteresis comparator and non-linear mapping successfully filter high-frequency noise, reducing the control signal flicker rate by 50.6% and converging the output signal standard deviation to 0.016. This research proves smooth, robust interactive feedback is achievable in complex stage environments using solely monocular vision, eliminating the need for expensive motion capture hardware and offering a viable solution for professional edge-computing stage interaction.