MAJEPPA is presented, a self-supervised framework to learn piano performance representations that span the full skill spectrum, from beginner practice sessions to virtuoso concert recordings, and a suite of downstream tasks spanning quality assessment, competition ranking, mistake and technique classification are introduced.
Abstract
We present MAJEPPA, a self-supervised framework to learn piano performance representations that span the full skill spectrum, from beginner practice sessions to virtuoso concert recordings. We curate the MAJEPPA dataset, comprising ~4,000 annotated recordings across six expertise levels and six recording contexts. We adapt a single pre-trained MIDI autoregressive model with a joint objective: next-token prediction learns score-conditioned performance generation at various skill levels, while InfoNCE and supervised contrastive losses align abstract score and performance representations in a joint embedding space. The proposed model both generates and understands performances in a unified framework. By introducing the EVPMR benchmark, a suite of downstream tasks spanning quality assessment, competition ranking, mistake and technique classification, we evaluate the learnt representations, demonstrating progress towards a real-world model for the piano performance space.
This survey traces attention from Bahdanau-Luong alignment through the Transformer and into vision architectures, and reviews fixed and learned sparse attention, linear attention, IO-aware exact algorithms including FlashAttention, and state-space alternatives including Mamba.
Results show that ARIMA is particularly efficient and effective on tasks involving harmonic, timing, and cross-performance retrieval, while remaining competitive with much larger baselines on other tasks.
A multi-layer feature fusion (MLF) adapter that aggregates information from all encoder layers before projecting them into the language model is proposed and shows that MICL does not emerge naturally in ALLMs, but can be effectively acquired through targeted contextual biasing training.
Xiaoyu Yang, Xuenan Xu, Wenyi Yu et al.· 0 citations
Lookahead Optimization for Rehearsal (LOR) is introduced, establishing a new and more robust paradigm for rehearsal-based OCIL and significantly out-performs state-of-the-art methods.
It is suggested that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling, in self-supervised fine-tuning.
Wangjin Zhou, Yizhou Zhang, Yichi Wang et al.· 0 citations
A strength-parity rule is proposed: an added member lowers the ensemble error only when it is both decorrelated from the current members and a near-peer of them in individual accuracy.