Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through succes...
Bing-Chen Yao, Hao-Bo Xu, Hao-Kun Lin et al.· 0 citations
This work introduces WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories, and builds Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's c...
Hao-Tian Zhang, Feng-Yuan Yu, Dezhi Luo et al.· 0 citations
This work introduces VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable, and identifies recurring failure modes of the prevalent VLM-as-a-judge paradigm.
Junhua Xu, Rui-Si Wang, Fanyi Pu et al.· 5 citations
OpenCoF, a framework comprising the OpenCoF-17K dataset, a reasoning video dataset spanning 11 task families, and Wan-CoF, a fine-tuned video model for studying whether diverse temporal supervision improves CoF behavior are introduced, suggesting that stronger video reasoning requires both broad temporal supervision an...