AWED, a word-level emission-delay metric defined relative to the acoustic end of each word, is introduced and post-train a DSM recognizer with GRPO under a reward that scores transcription accuracy and measured delay jointly, advancing the accuracy--latency Pareto frontier of streaming ASR without architectural change.
Abstract
Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $\tau$ that bounds the decoder's lookahead. We show that $\tau$ is a poor proxy for user-perceived latency, and that the alignment-based supervision of DSM leaves latency on the table: the same forced-aligned transcript is used at every $\tau$, forcing the model to withhold words it could already commit. We introduce AWED, a word-level emission-delay metric defined relative to the acoustic end of each word, and post-train a DSM recognizer with GRPO under a reward that scores transcription accuracy and measured delay jointly. Trained at a single operating point ($\tau=6$ frames), our model dominates both its supervised fine-tuning initialization and the Voxtral Realtime backbone across all evaluated lookahead budgets: it cuts WER by 30.8\% relative at an 80\,ms structural delay, and by 5.7\% relative at 480\,ms while lowering median AWED from 1.17\,s to 1.04\,s. Latency-rewarded post-training thus advances the accuracy--latency Pareto frontier of streaming ASR without architectural change.
X2Streaming-ASR is proposed, which decomposes streaming recognition into when to commit and what to commit, and achieves the best streaming CER among the evaluated systems on AISHELL-1 and AISHELL-3 with substantially lower latency.
Zhi-Wei Lin, Kaiqi Fu, Rime Wen et al.· 0 citations
This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a...
In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition, where demonstrations merely cue pre-trained capabilities, rather than Task Learning, where a genuinely new input-label mapping must be inf...
Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al.· 0 citations
The resulting training-free pipeline, RECAST, consistently outperforms recent frame-selection baselines without modifying the ITM encoder, and preserves the per-frame matching cost of a standard single-query baseline.
S. Han, Thang Vu, Junyeong Kim· IEEE Access· 0 citations
This work studies how to compress a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share.
Prasanth Yadla, Mohammad Samragh, Dongseong Hwang et al.· 0 citations
In this paper, we propose an alignment-path distillation framework for streaming automatic speech recognition (ASR) with large language models (LLMs). Interleaved streaming ASR-LLMs use forced alignments (FA) from alignment models, such as those trained with connectionist temporal classification (CTC), to construct spe...
Yan-Song Jia, Kai-Wei Huang, Junjie Chen et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.