Speech-to-LLM systems often connect a frozen speech encoder to a frozen large language model (LLM) through a small trainable bridge. The bridge is usually treated as plumbing, but it in fact defines the geometry of the speech-to-LLM interface, and the pretraining objective decides whether that interface provides a reus...
Xin-Nian Zhao, Chia-Hua Wu, Pu Wang et al.· 0 citations
End-to-end attention-based speech recognition is accurate offline but hard to stream: outputs can depend on future audio, and a little future context per layer makes the lookahead grow with the number of layers. We address this with two mechanisms. A bounded-lookahead chunk encoder caps every chunk's future receptive f...
Yi-Chen Jia, Bastiaan Tamm, H. Van Hamme· 0 citations
Automatic speech recognition models suffer from catastrophic forgetting when adapted to new domains, accents, or downstream tasks. This problem becomes increasingly important with the growing use of speech foundation models, where adaptation should be both memory-efficient and safe, preserving the broad capabilities le...
Steven Vander Eeckt, H. Van Hamme· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.