Preprint
Jul 2026
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and text-based reasoning.
Xu Wang, Kaixiang Yao, Miao Pan et al.
· 1 citation