Skip to content

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

Sep 2026 · 0 citations · 18 references
Engineering Computer Science

TL;DR

A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction, and introduces a fixed-receptive-field convolutional encoder that reduces the respective prediction errors.

Abstract

Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction. Guided by this finding, we introduce a fixed-receptive-field convolutional encoder that reduces the respective prediction errors by 36.0%, 17.3%, and 3.4%. We further show that directly transferring image-domain gradient-variance supervision restores fine-scale variation but degrades predicted quality, motivating a Mel-specific formulation with axis-specific gradients, overlapping local statistics, and log-domain variance matching. GrainSpeech contains only 264.8K parameters and achieves 17.9x real-time Mel generation on a microcontroller (MCU), while attaining UTMOS scores comparable to substantially larger models with less than 1.5% of their parameters. Source code and demos are available at https://github.com/lab-emi/GrainSpeech.

View source

Similar papers

Preprint Sep 2026

HAMMER: Harmonic-Aware Parallel Context Modeling and Discriminator-Free Perceptual Optimization for Speech Enhancement

Recent speech enhancement systems combine self-attention and Mamba to capture global interactions and long-range dependencies. Yet these hybrids usually operate as sequence mixers and do not explicitly exploit harmonic periodicity, a strong cue for preserving voiced speech under noise. Perceptual optimization poses ano...

Shang-Fu Chen, Szu-Wei Fu, Sung-Feng Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries

Results establish MADS (Multi-view Acoustic Descriptor Set) not merely as a competitive standalone descriptor set, but as the foundational descriptor layer of a broader acoustically grounded representation program for future frame-level and deep-learning-compatible audio modeling.

Utsab Ghosh, Roshni Chakraborty · 0 citations
#artificial intelligence Preprint Sep 2026

GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations a...

G. Botté, Séverin Baroudi, Samir Sadok et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection

CRAF, a cross-view residual-aware fusion framework that uses ALLM-guided cross-view attention to enrich SSL representations and adopts ALLM as a high-level reference to separate ALLM-explainable information from complementary SSL residual information, demonstrates robustness to unseen spoofing attacks.

M. Phan, Khalid Zaman, C. Mawalim et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.