Skip to content

Author

Ziyang Jiang

We have 2 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Sep 2026

Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities

Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems conditioned on enrollment speech have progressed substantially,...

Xin-Yuan Qian, Yang Zhou, Zi-Yang Jiang et al. · 0 citations
Preprint Aug 2026

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

SpeechAgent-R, an audio agent that coordinates its intrinsic multimodal understanding with external skills and tools, is developed and HIU-Bench is introduced to jointly evaluate task performance, interaction quality and generalization to diverse task settings.

Yuwen Wang, Tian-Hao Zhang, Ming Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.