In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtime interaction with an asynchronous agent loop. In contrast to conventional turn based systems, Gander continuously processes streaming user inputs, enabling full-duplex interactio...
Orantqing, Shengpeng Ji, Jun-Long Tong et al.· 1 citation
Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying part...
Wen-Xu Jia, Xi-Ze Cheng, Zi-Han Zhang et al.· 0 citations
While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation...
Dongjie Fu, Di Cao, Xize Cheng et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.