Skip to content

Author

Xiangang Li

17 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Qwen-Audio-Agent Technical Report

We present Qwen-Audio-Agent, a harness that combines full-duplex voice interaction with asynchronous task execution through a foreground-background architecture. A Frontend Agent manages dialogue and selects between direct tool use and delegation, while a Backend Agent carries out delegated tasks in a separate context....

Chong Deng, Yunjie Ji, Yu-Xiang Kong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

agentic-ger: terminology recovery in long-form speech using global context

Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agenti...

Yan-Qiao Zhu, Wu-Peng Wang, Zhifu Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Dis...

Lu-Jia Bao, Qian Chen, Luyao Cheng et al. · 1 citation
Preprint Sep 2026

The Evolving Bottleneck in Speech Generation: Interface Co-design and Staged Alignment from CosyVoice to Qwen-Audio-3.0-TTS

Speech synthesis systems are commonly narrated as a sequence of larger models, better tokenizers, and broader data. This technical retrospective offers a different account of the CosyVoice lineage, from CosyVoice through CosyVoice 2 and CosyVoice 3 to Qwen-Audio-3.0-TTS: progress came from repeatedly relocating the sys...

Qian Chen, Xiangang Li, Xiang Lv et al. · 0 citations
Preprint Sep 2026

Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark

Conversational context provides semantic and acoustic cues across turns for automatic speech recognition (ASR), but relying on historical transcripts can propagate recognition errors and discard pronunciation and speaker information. We present a multimodal conversational-context framework for LLM-based ASR that integr...

Long-Hao Li, Jian Tang, Yu-Xiang Kong et al. · 0 citations
Conference Open access Sep 2026

FunCineForge: A Unified Dataset Pipeline and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

Movie dubbing is the task of synthesizing speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. However, existing methods face two major limitations: (1) high-quality multimodal dubbing datasets are limited in scale...

Jia-Xuan Liu, Yang Xiang, Han Zhao et al. · 0 citations
Preprint Sep 2026

TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models

Large audio language models (LALMs) can describe what is heard, but their ability to localize when queried content occurs remains less systematically evaluated. We present TAG-Bench, a benchmark for temporal audio grounding in which a model returns every time interval that matches a natural-language query. TAG-Bench co...

Yu-Hang Dai, Xin Shu, Zeng-Xi Li et al. · 0 citations
Preprint Sep 2026

Semantic Refinement of Universal Audio Representations through Audio-Description Alignment

Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of...

Le-Jun Min, Jun-Yu Dai, Rui-Chen Zheng et al. · 0 citations
Review Jul 2026

Qwen-Audio-3.0-Gen-Preview Technical Report

Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion Transformer (DiT) and a shared variat...

Jun-Yu Dai, Xiao-Yue Duan, Xin-Yu Fan et al. · 1 citation
Jul 2026

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text lets the model copy instead of listen, so...

Peng-Fei Zhang, Biao Tian, Tianxin Xie et al. · 0 citations
#machine learning Preprint Sep 2026

The Platonic brain bridge hypothesis: human brain networks as an architectural prior for multimodal large language models

Multimodal large language models predict brain activity, but brain alignment has been a measurement, not a design tool. We propose the Platonic brain bridge hypothesis: omni models, multimodal large language models that process video, audio and text jointly, converge on brain-like representations usable in both directi...

Peng-Fei Zhang, Biao Tian, Xian-Gang Li et al. · 0 citations
#natural language process... Preprint Sep 2026

Qwen-Audio-3.0-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility r...

Chuan-Meng Bian, Da-Ren Chen, Pei-Xin Chen et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.