Vision-Language Models (VLMs) have advanced rapidly in static visual understanding, yet remain unreliable when judging how an egocentric task is progressing. Given a task instruction and two visual observations, a model should determine which state is closer to the goal by analyzing task-relevant object configurations...
Xiao-Da Yang, Can Wang, Yu-Xiang Liu et al.· 0 citations
In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtime interaction with an asynchronous agent loop. In contrast to conventional turn based systems, Gander continuously processes streaming user inputs, enabling full-duplex interactio...
Orantqing, Shengpeng Ji, Jun-Long Tong et al.· 1 citation
A rubric-based audio-grounded evaluation framework that verifies event realization, acoustic attributes, and speech content through fine-grained semantic criteria, and human validation demonstrates that the benchmark achieves stronger alignment with human semantic judgments than conventional global similarity metrics.
Jin-Ting Wang, Yuguang Yang, Shengyu Li et al.· 0 citations
Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA, a distillation and inference framework for accelerating a 19B-parameter joint...
Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng et al.· 0 citations
Doc-REFRAG, a question-guided framework that compresses visual tokens into coarse chunks and selectively expands question-relevant ones via a lightweight RL-based selector, is proposed, achieving state-of-the-art accuracy with significantly lower inference latency.
Ruofan Hu, Sheng Xu, Min-Jie Hong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.