Skip to content

Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR

Sep 2026 · 0 citations · 19 references
Computer Science

TL;DR

AWED, a word-level emission-delay metric defined relative to the acoustic end of each word, is introduced and post-train a DSM recognizer with GRPO under a reward that scores transcription accuracy and measured delay jointly, advancing the accuracy--latency Pareto frontier of streaming ASR without architectural change.

Abstract

Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $\tau$ that bounds the decoder's lookahead. We show that $\tau$ is a poor proxy for user-perceived latency, and that the alignment-based supervision of DSM leaves latency on the table: the same forced-aligned transcript is used at every $\tau$, forcing the model to withhold words it could already commit. We introduce AWED, a word-level emission-delay metric defined relative to the acoustic end of each word, and post-train a DSM recognizer with GRPO under a reward that scores transcription accuracy and measured delay jointly. Trained at a single operating point ($\tau=6$ frames), our model dominates both its supervised fine-tuning initialization and the Voxtral Realtime backbone across all evaluated lookahead budgets: it cuts WER by 30.8\% relative at an 80\,ms structural delay, and by 5.7\% relative at 480\,ms while lowering median AWED from 1.17\,s to 1.04\,s. Latency-rewarded post-training thus advances the accuracy--latency Pareto frontier of streaming ASR without architectural change.

View source

Similar papers

#natural language process... Preprint Sep 2026

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a...

H. Le, L. Nguyen, Minh Tri Dao · 1 citation
Preprint Sep 2026

AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning

In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition, where demonstrations merely cue pre-trained capabilities, rather than Task Learning, where a genuinely new input-label mapping must be inf...

Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al. · 0 citations
Open access 2026

Speaking ITM’s Language: Query Reformulation and Temporal-MMR for Long-Video QA

The resulting training-free pipeline, RECAST, consistently outperforms recent frame-selection baselines without modifying the ITM encoder, and preserves the per-frame matching cost of a standard single-query baseline.

S. Han, Thang Vu, Junyeong Kim · 0 citations
#small language model Preprint Sep 2026

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

This work studies how to compress a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share.

Prasanth Yadla, Mohammad Samragh, Dongseong Hwang et al. · 0 citations
Preprint Sep 2026

Alignment-Path Distillation from Non-streaming ASR-LLMs for Streaming Speech Recognition

In this paper, we propose an alignment-path distillation framework for streaming automatic speech recognition (ASR) with large language models (LLMs). Interleaved streaming ASR-LLMs use forced alignments (FA) from alignment models, such as those trained with connectionist temporal classification (CTC), to construct spe...

Yan-Song Jia, Kai-Wei Huang, Junjie Chen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.