In-Context Adaptation of Encoder-Decoder Models in Speech Recognition
This work studies two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures and suggests that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.