Visual Language Models (VLMs) have achieved remarkable success across diverse tasks, yet they struggle with high-resolution inputs where critical information resides in small regions or detailed, cluttered scenes. While several approaches address this limitation, a systematic understanding of why models fail at high re...
This work introduces a human-aligned evaluation framework for text-to-SVG generation, and develops two complementary evaluators: CLIP scorers adapted to vector graphics and then aligned to human preferences, for fast large-scale evaluation and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcem...
Marco Cipriano, Leonardo Zini, Alexandra Schild et al.· 0 citations
This work proposes a novel test-time alignment approach that leverages trajectory-guided structured sampling for dynamic inference-time refinement, achieving better alignment with visual grounding and ensuring logical consistency, and demonstrates that this approach significantly improves accuracy without incurring pro...
Tian-Bao Jiang, Wei-Cong Ni, Gerard de Melo et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.