Skip to content

Author

Zhizheng Wu

9 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

GenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages

Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain...

Li Wang, Kun-Yu Feng, Wan Lin et al. · 0 citations
#machine learning Preprint Oct 2026

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic...

Yu-Xiang Wang, Kun-Yu Feng, Yuan-Cheng Wang et al. · 0 citations
Preprint Sep 2026

Interpreting and Evaluating Dynamic-Rate Speech Codec Boundaries

Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful for neural speech reconstruction. This work combines bounda...

Han Wang, Jia-Qi Li, Ying Shen et al. · 0 citations
#machine learning Preprint Sep 2026

EvoAudio: Recursive Self-Improvement for Audio Understanding

Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-im...

Yu-Xiang Wang, Sheng-Bo Cai, Ying Shen et al. · 0 citations
#natural language process... Preprint Sep 2026

Spoken Language Models that Think Aloud

While Chain-of-Thought (CoT) reasoning has improved the capability of language models, directly applying it to Spoken Language Models (SLMs) may introduce long silent intervals under the serial"think-then-speak"paradigm, disrupting real-time spoken interaction. To address this issue, we propose an asynchronous think-al...

Junyi Ao, Kainan Peng, Mingbo Ma et al. · 0 citations
Jul 2026

Teffic-Audio: Tell Fact from Fiction

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and tra...

Wan Lin, Li Wang, Jindong Wang et al. · 1 citation

ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models

ParaBridge is proposed, an on-policy self-distillation method that turns a brittle inference-time scaffold into stable model behavior and generalizes to unseen paralinguistic cues, transfers from safety-oriented training to empathy-oriented dialogue, and works on a different SLM backbone.

Yuxiang Wang, Qin-Ke Ni, Sheng-Bo Cai et al. · 2 citations
#machine learning Preprint Sep 2026

RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory

RecurTrace introduces Loop Memory Attention, which lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone.

Yu-Xiang Wang, Kun-Yu Feng, Ying-Da Shen et al. · 3 citations

VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models

The VoxPrivacy benchmark, the large-scale training set, the fine-tuned model to foster the development of safer and more context-aware SLMs, and a viable path forward are demonstrated: by fine-tuning on a new 4,000-hour training set, improve privacy-preserving abilities while maintaining robustness.

Yu-Xiang Wang, Hongyu Liu, De-Kun Chen et al. · 5 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.