Skip to content

Author

Mu-Zhi Zhu

We have 3 of 34 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external c...

Haolin He, Yun-Fei Chu, Qi Chen et al. · 0 citations
Preprint Sep 2026

Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes

Metric reasoning is a critical and challenging task for Vision Language Models (VLMs), playing a pivotal role in embodied AI tasks such as robotic manipulation and autonomous navigation. However, current spatial reasoning remains bottlenecked by rigid pixel-level supervision; such localized optimization often compromis...

Yu-Ling Xi, Hao-Kai Zhang, Mu-Zhi Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external c...

Haolin He, Yunfei Chu, Qi Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.