Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation amo...
Chengqian Ma, Wen-Hao Feng, Wei-Xuan Jin et al.· 0 citations
WaveZip is proposed, a joint signal-frequency-domain framework for efficient video inference that requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency.
Yuhui Zeng, Wang Chen, Jin-Fa Huang et al.· arXiv.org· 1 citation
Modality Subspace Activation (MSA) is proposed, a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths and dynamically balances modal projections in the last hidden state, effectively restoring CMS across benchmarks.
Hongbo Jiang, Jie Li, Yunhang Shen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.