Skip to content
Preprint

SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations

Sep 2026 · 0 citations · 26 references
Engineering Computer Science

Abstract

Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned feature aggregation consolidates complementary linguistic and paralinguistic information across a frozen encoder into discrete targets for masked prediction, while a coarse-to-fine objective combines log-Mel reconstruction with residual flow matching to preserve spectral structure and fine-grained acoustic variation. Experiments on SUPERB and speech resynthesis show that SPEAR-Gen maintains strong understanding performance while substantially improving resynthesis quality and speaker preservation. These results demonstrate that a single speech representation can effectively support both understanding and generation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.