Skip to content
Conference

Style-aware data augmentation for deep learning on symbolic music

Jul 2026 · International Conference on Image, Video and Signal Processing · Vol 14268, pp. 1426811 - 1426811-13 · 0 citations · 22 references
Engineering

Abstract

This study proposes a style-aware data augmentation framework that combines rule-based design with statistical constraints, applied to small-scale, highly constrained Jiangnan symbolic music generation tasks. By leveraging YNote representation, fixed rhythmic frameworks, and Markov-style local transition statistics, we systematically expand the training data while maintaining musical structural plausibility. Using the augmented data, we fine-tune a GPT-2 model to analyze how different training data scales affect generation behavior. Experimental results show that Bilingual Evaluation Understudy (BLEU) -based reference-overlap metrics exhibit only minor fluctuations across different training scales and are insufficient to directly reflect style improvement. In contrast, Kullback-Leibler (KL) divergence and bigram statistics effectively characterize the style consistency of the generated set in terms of overall distribution proximity and local transition plausibility. Further analysis indicates that a medium-scale training set (approximately 3,000–12,000 samples) achieves the best balance between distribution consistency and transition coverage, whereas excessive augmentation may lead to distribution calibration drift. Overall, the study demonstrates that data augmentation has a positive but non-monotonic effect on style consistency, highlighting the need for carefully designed augmentation strategies in highly constrained symbolic music generation tasks.

View source