State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models
Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approaches rely on globally compressed visual tokens, where fine-grained details are entangled within a single representation and repeatedly accesse...