Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approaches rely on globally compressed visual tokens, where fine-grained details are entangled within a single representation and repeatedly accesse...
Ming-Xu Chai, Chen-Yu Liu, Zi-Yu Shen et al.· 0 citations
The Prefix-Adaptive Block Diffusion Model (PA-BDM) is proposed, which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit.
Ming-Xu Chai, Zi-Yu Shen, Chen-Yu Liu et al.· arXiv.org· 0 citations
HIEVI-RAG is introduced, a hierarchical, evidence-driven multimodal RAG framework for closed-domain document understanding that significantly outperforms existing open-source baselines and exceeds the strongest reported baseline by an average of 8.05% in accuracy.
Junyu Xiong, Yonghui Wang, Rongjian Gu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.