Skip to content

In-Context Adaptation of Encoder-Decoder Models in Speech Recognition

Sep 2026 · 0 citations · 48 references
Computer Science Engineering

TL;DR

This work studies two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures and suggests that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.

Abstract

In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, when providing interleaved speech-text demonstrations. In this work, we ask whether in-context adaptation is an inherent ability for all encoder-decoder models. We study two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures. We find that all tested models are able to perform in-context adaptation out of the box, achieving up to 30% relative improvement in the oracle experiments and up to 23% using first-pass hypotheses. Through controlled experiments on three English datasets, we show that lexical and speaker information both contribute to successful adaptation. While interleaved demonstration is effective in certain cases, collated demonstration brings consistent adaptation across the board. Our results suggest that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.

View source

Similar papers

Models are Zero-Shot Text

Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.

Unknown authors · 0 citations
#machine learning Preprint Sep 2026

Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation

Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To ad...

Filip Tăşădan, Elena Tomanová, Ondrej Lopuch et al. · 0 citations
Open access Sep 2026

Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods

Neural text-to-speech (TTS) systems can synthesize highly natural speech. A key capability is speaker adaptation, which enables speech synthesis that matches a target speaker’s voice characteristics, such as timbre, pitch, and prosody. Traditional neural approaches require extensive speaker-specific data and full retra...

Kishor Kayyar Lakshminarayana, Frank Zalkow, Christian Dittmar et al. · 0 citations
Preprint Aug 2026

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al. · 0 citations
Preprint Sep 2026

Beyond Encoder Fusion: Multi-View Discrete Token Augmentation for LLM-Based ASR

Discrete speech tokens provide a compact interface between speech encoders and large language models for automatic speech recognition, but single-tokenization systems remain sensitive to the chosen encoder. We propose multi-view discrete token augmentation, a simple strategy that augments each training utterance by gen...

Paul Moïse Gangbadja, Mickael Rouvier, Fabrice Lefèvre · 0 citations
#natural language process... Preprint Sep 2026

Merging the Knowledge of LLMs for Automatic Speech Recognition

This method integrates the LMs directly into the parameters of an LLM-based ASR model, requiring no additional computational cost at inference, and consistently improved the ASR performance in the target domains, without degrading inference speed or memory footprint.

Hayato Futami, Tatsuya Kawahara · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.