Preprint
Jul 2026
Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering
The findings demonstrate that while large pretrained VLMs possess strong zero-shot baselines for structured layouts, their performance strongly decreases on visually complex layouts of infographics and slides, and that visual understanding is the main bottleneck for DocVQA, not a lack of knowledge from the VLMs.
Miguel Lopez-Duran, Elena Marrero, Julian Fiérrez et al.
· 2 citations