The Document Question Answering (DocQA) task necessitates the synergistic interpretation of visual and textual information embedded within documents. Although Retrieval-Augmented Generation (RAG) has enhanced the capabilities of Large Vision-Language Models (LVLMs), existing approaches still encounter significant bottl...
Jia-Yuan Wang, Jie Lian, Fu Zhao et al.· Proceedings of the 32nd ACM...· 0 citations
Center-guided spectral diffusion is proposed, which replaces traditional alignment with generative modeling and outperforms state-of-the-art methods on several datasets and alleviates the noise amplification problem commonly found in traditional alignment methods.
Jia-Yuan Wang, Jie Lian, Yong-Quan Shi et al.· Proceedings of the Thirty-Fi...· 0 citations
Unsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. The...
Xuan Tan, Qi-Xian Zhang, Ding Qi et al.· IEEE Transactions on Image P...· 0 citations
This work introduces MMHRAG, a novel Multi-Modal Hierarchical Retrieval-Augmented Generation framework that achieves cross-modal interaction on DocQA for the first time, and designs a Summarizing Agent to resolve logical conflicts, information redundancy, and granularity discrepancies among retrieved cross-modal eviden...
Jiayuan Wang, Jie Lian, Fu Zhao et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.