Skip to content

Voice Memory for Agentic Speech Recognition

Jul 2026 · arXiv.org · Vol abs/2607.26410 · 0 citations · 62 references
Computer Science Engineering

TL;DR

Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best, and transfers across corrector families and adds zero parameters to the inference path.

Abstract

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

agentic-ger: terminology recovery in long-form speech using global context

Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agenti...

Yan-Qiao Zhu, Wu-Peng Wang, Zhifu Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VoiceLongMemEval: Do Assistants Remember How You Sounded?

VoiceLongMemEval (VLME) benchmark, where every answer depends on paralinguistic metadata attached to conversational turns, which is otherwise unrecoverable from the words alone, is presented.

Ramit Pahwa, Parivesh Priye, Apoorva Beedu · 1 citation
#natural language process... Preprint Sep 2026

Voice of Reason: Reinforcement Learning for Spoken Math

Speech language models enable richer spoken interactions between humans and machines than cascaded systems, allowing access to paralinguistic information and lower latency. However, their accuracy on mathematical reasoning benchmarks has lagged behind those of text models. Reinforcement learning (RL) with verifiable re...

Timothée Weisselberger, Edouard Graves, Alexandre Défossez · 0 citations
#artificial intelligence Preprint Sep 2026

Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents

Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Each ite...

Yan-Jie Zhang, Nan-Chen Hu, Yu-Shi Sun · 0 citations
#artificial intelligence Preprint Sep 2026

RetroThinker: Enabling Retrospective Thinking in Speech LLMs

Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imp...

Yi-Jen Shih, Pu-Yuan Peng, Abdel-rahman Mohamed et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.