Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-language models often treat motion as an auxiliary modality of a language model, leading to text-dominated representations and limited cross-mod...
Guo-Cun Wang, Kenkun Liu, Guo-Rui Song et al.· 0 citations
Large vision-language models (VLMs) are increasingly deployed in safety-critical settings, yet existing visual jailbreak research has focused almost exclusively on autoregressive architectures, leaving an important emerging family unstudied: multimodal discrete diffusion vision-language models (dVLMs). We identify a vu...
Guo-Rui Song, Run-Qing Tang, Jing-Ye Zhang et al.· 0 citations
A Video Compression framework built upon a foundational flow model that enables the compressor to harness generative video flow priors effectively, which reduces bit consumption by 58\% and achieves one-step decoding and reconstructions with high perceptual fidelity.
Yichong Xia, Qin-Hong Wu, Bin Chen et al.· 1 citation· ⚡1
R3G, a modular Reasoning-Retrieval-Reranking framework, first produces a brief reasoning plan that specifies the required visual cues, then adopts a two-stage strategy, with coarse retrieval followed by fine-grained reranking, to select evidence images.
Zhuo Chen, Zhengxian Wu, Zi-Rui Liao et al.· IEEE International Conferenc...· 4 citations
KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.
Hangrui Xu, Zheng-Xian Wu, Yu Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.