Skip to content
Preprint

Selective Lookahead for Attention-Based Streaming ASR

Sep 2026 · 0 citations · 51 references
Engineering

Abstract

End-to-end attention-based speech recognition is accurate offline but hard to stream: outputs can depend on future audio, and a little future context per layer makes the lookahead grow with the number of layers. We address this with two mechanisms. A bounded-lookahead chunk encoder caps every chunk's future receptive field at a constant number of chunks, independent of the number of layers, via one age-selection rule shared by self-attention and the depthwise convolution. On this encoder, dynamic future-chunk decoding lets a per-token trigger commit a token or wait and re-decode it; we propose a learned trigger as the general mechanism, with a simple confidence threshold as an effective fallback. On full LibriSpeech test-clean the dynamic system matches the best static-lookahead accuracy (6.5%) at a median latency of 306 ms versus 860 ms for one-chunk static lookahead, and a wait budget bounds the deferral tail below the static baseline's 90th percentile at 0.1 points more WER.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.