An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first gen...
Chengqian Ma, Wei Tao, Hao-Yu Zhang et al.· 0 citations
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning...
Tianchi Liu, Ze-Yang Song, Tian-Rui Wang et al.· 3 citations
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whe...
Jing-Di Lei, Junxian Li, Di Zhang et al.· 0 citations
Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied...
Yun-Cheng Guo, Zhan-Qiu Zhang, Yiwen Guo et al.· 0 citations
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, but a spoken answer still leaves the agent visually absent. We introduce \textbf{Ex-Omni-2D}, a framework that answers a multimodal query with coordinated text, personalized speech, and reference-conditioned video. The dialogue m...
Haoyu Zhang, Zhipeng Li, Xiaoying Tang et al.· 0 citations
This work constructs two human-verified benchmarks, VRQABench for controllable spatial lookahead and OpenWorldQA for open-domain physical prediction, and proposes Privileged-Future On-Policy Self-Distillation (PF-OPSD), a model that outperforms baseline by 10.6% and 10.9% on VRQABench and OpenWorldQA, respectively, whi...
Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair progress implicit. This causes inefficient repair targeting, slow conso...
Lekang Jiang, Bohan Tang, Stephan M. Goetz et al.· 0 citations
MemDefrag, a training-free and model-agnostic framework that uses a middle-layer tracing signal to conduct memory defragmentation (rank, reorder, and filter memories), and applies an informativeness-guided proportional forgetting mechanism once capacity is exceeded, is proposed.
Ruiyi Yan, Zhuoyuan Mao, Yiwen Guo· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.