Skip to content

Author

Haoqian Wang

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World

Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-language models often treat motion as an auxiliary modality of a language model, leading to text-dominated representations and limited cross-mod...

Guo-Cun Wang, Kenkun Liu, Guo-Rui Song et al. · 0 citations
Preprint Sep 2026

DIVA: Exploiting Cross-Step Conditional Propagation for Visual Jailbreaks in Discrete Diffusion Vision-Language Models

Large vision-language models (VLMs) are increasingly deployed in safety-critical settings, yet existing visual jailbreak research has focused almost exclusively on autoregressive architectures, leaving an important emerging family unstudied: multimodal discrete diffusion vision-language models (dVLMs). We identify a vu...

Guo-Rui Song, Run-Qing Tang, Jing-Ye Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VoRTeC: Taming Foundation Flow for One-step Real time Video Compression

A Video Compression framework built upon a foundational flow model that enables the compressor to harness generative video flow priors effectively, which reduces bit consumption by 58\% and achieves one-step decoding and reconstructions with high perceptual fidelity.

Yichong Xia, Qin-Hong Wu, Bin Chen et al. · 1 citation · ⚡1
#artificial intelligence Conference Jan 2026

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation

R3G, a modular Reasoning-Retrieval-Reranking framework, first produces a brief reasoning plan that specifies the required visual cues, then adopts a two-stage strategy, with coarse retrieval followed by fine-grained reranking, to select evidence images.

Zhuo Chen, Zhengxian Wu, Zi-Rui Liao et al. · 4 citations
Preprint Aug 2026

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.

Hangrui Xu, Zheng-Xian Wu, Yu Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.