Skip to content
Preprint

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

Sep 2026 · 2 citations · 60 references
Computer Science Engineering

TL;DR

Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction, which leads the compared models on MMAU, Llama Questions, and Speech CMMLU.

Abstract

Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events. A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue. Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows. Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.

View source

Similar papers

Preprint Sep 2026

Multimodal Duplex Interaction Agent

Gander is presented, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtime interaction with an asynchronous agent loop and is released together with its models, code, and data to facilitate further research and development in the community.

Orantqing, Shengpeng Ji, Jun-Long Tong et al. · 2 citations
Preprint Sep 2026

StepAudio 3 Realtime Technical Report

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret user intent, while Seamless Duplex model...

Bin Lin, Bo Zhao, Bo-Yang Zhang et al. · 3 citations
Preprint Sep 2026

Qwen-Audio-Agent Technical Report

We present Qwen-Audio-Agent, a harness that combines full-duplex voice interaction with asynchronous task execution through a foreground-background architecture. A Frontend Agent manages dialogue and selects between direct tool use and delegation, while a Backend Agent carries out delegated tasks in a separate context....

Chong Deng, Yunjie Ji, Yu-Xiang Kong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Dis...

Lu-Jia Bao, Qian Chen, Luyao Cheng et al. · 1 citation
#natural language process... Preprint Sep 2026

SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents

SALMONN-duo is proposed, an adaptive dual-system voice agent inspired by dual-process theories of cognition that separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM with a powerful asynchronous slow-thinking LLM agent.

Wen-Yi Yu, Si-Yin Wang, T. Chiba et al. · 0 citations
#artificial intelligence Preprint Sep 2026

APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction

Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completio...

Puneet Mathur, Dinesh Manocha · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.