Style-aware data augmentation for deep learning on symbolic music
This study proposes a style-aware data augmentation framework that combines rule-based design with statistical constraints, applied to small-scale, highly constrained Jiangnan symbolic music generation tasks. By leveraging YNote representation, fixed rhythmic frameworks, and Markov-style local transition statistics, we systematically expand the training data while maintaining musical structural plausibility. Using the augmented data, we fine-tune a GPT-2 model to analyze how different training data scales affect generation behavior. Experimental results show that Bilingual Evaluation Understudy (BLEU) -based reference-overlap metrics exhibit only minor fluctuations across different training scales and are insufficient to directly reflect style improvement. In contrast, Kullback-Leibler (KL) divergence and bigram statistics effectively characterize the style consistency of the generated set in terms of overall distribution proximity and local transition plausibility. Further analysis indicates that a medium-scale training set (approximately 3,000–12,000 samples) achieves the best balance between distribution consistency and transition coverage, whereas excessive augmentation may lead to distribution calibration drift. Overall, the study demonstrates that data augmentation has a positive but non-monotonic effect on style consistency, highlighting the need for carefully designed augmentation strategies in highly constrained symbolic music generation tasks.