Research on Animation Generation Method of Virtual Digital Human Based on Motion Capture and AI-Driven Technology
Abstract
Virtual digital human animation generation faces persistent challenges in motion fidelity, facial expressiveness, and whole-body coordination. A hybrid framework integrating optical multi-camera motion capture with AI-driven generation models is proposed to address these limitations. The system employs a 16-camera optical capture array calibrated to sub-millimeter precision, combined with a Bidirectional Long Short-Term Memory (BiLSTM) network for motion sequence generation and an audio-semantic fusion module for facial animation synthesis. Experimental results demonstrate that the proposed framework achieves a mean joint position error (MPJPE) of 18.3 mm, a Fréchet Inception Distance (FID) score of 6.72 in motion generation quality, and a lip-synchronization accuracy of 94.6%. Compared with existing single-modal methods, the proposed approach reduces temporal jitter by 37.4% and improves rendering frame stability to 58.2 FPS under real-time conditions. The results confirm that the combined capture-and-generation pipeline effectively improves animation realism and computational efficiency, offering a scalable solution for virtual human deployment across entertainment, education, and interactive media domains.