Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer-grained supervision primarily rely on textual de...
Chunyi Peng, Zhipeng Xu, Y. Yukun et al.· 0 citations
Single domain generalization (SDG) aims to learn a model from one labeled source domain that generalizes to unseen target domains. A common strategy is to enrich the source distribution with augmented or generated samples, and recent text-to-image (T2I) diffusion models provide a strong generative prior for this purpos...
Zhi-Peng Xu, De Cheng, Xinyang Jiang et al.· 0 citations
A paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context shows that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance...
Zhi-Peng Xu, Jia-Hao Lu, Yi-Ning Zheng et al.· 4 citations
SAYRE is presented, a scene-aware document synthesis framework for generating scalable KIE training data without hand-crafted template design, and error analysis shows that synthesized training reduces field-level errors by improving schema-aware extraction over dense tables, business identifiers, and contract clauses.
Zhipeng Xu, Zulong Chen, Qing Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.