Intelligent Interpretation of Traditional Folk Tales through Vocalization: a Generative Model Based on Emotion-Driven and Lip-Syncing
Abstract
This study presents an integrated multi-modal signal generation framework for intelligent vocalization of traditional folk tales, ensuring emotion-driven speech synthesis with precise lip synchronization. The proposed system models continuous narrative emotion trajectories and embeds them into phoneme-level acoustic generation, using differentiable temporal alignment to synchronize audio and visual outputs. Multi-channel signals—including acoustic features and lip motion vectors— are processed through recursive smoothing and alignment matrices to maintain temporal consistency across modalities. Experimental results demonstrate median phoneme-duration deviations of 19.4–22.8 ms and lip-sync errors within 0.031–0.034, indicating stable audio-visual synchronization under complex narrative rhythms. By interpreting the system as a multi-channel signal processing and closed-loop control framework, the method mirrors principles of precision-engineered communication and control systems, providing insights into temporal alignment, cross-modal consistency, and real-time feedback optimization. This approach offers an engineering-oriented perspective for the development of advanced audio-visual generation, multi-sensor signal integration, and temporally coherent generative models.