Skip to content

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

Sep 2026 · 0 citations · 45 references
Computer Science

TL;DR

This work studies how to compress a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share.

Abstract

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share. We train only the student encoder to regress the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8x compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.

View source

Similar papers

#machine learning Preprint Sep 2026

SAGE: Semantic Audio Generative Encoder

Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of thes...

Francesco Brigante, Luca Cerovaz, Davide Marincione et al. · 0 citations
#artificial intelligence Preprint Sep 2026

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short beha...

Hao-Jun Zhang, Yi Zou, Min Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LatentPress: Context Compression Beyond Text and Vision

LatentPress is introduced, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no text reconstruction at inference.

Zheng-Ze Zhou, Hejian Sang · 1 citation
#artificial intelligence Preprint Sep 2026

Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation

Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selecte...

Xiao-Yu Wu, Wei-Hang Guo, Yi-Fei Wang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.