Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-policy RLVR algorithms face a critical optimization bottleneck in preserving and reinforcing visually grounded reasoning behaviors: valuable visually-...
Xin-Xin Song, Si-Yuan Li, Tingxiong Xiao et al.· 0 citations
ExBind is designed for controlled diagnosis rather than population-scale ranking or end-to-end editing evaluation, and samples representation-independent latent binding instances and compiles them into SVG, DOM, canvas, tree, graph, and table cases with deterministic mappings to executable references.
Zi-Qian Wang, Yuxiao Cheng, Tingxiong Xiao et al.· 0 citations
CARE is proposed, which grounds every rubric evolution step in a high-quality anchor response generated by a frontier model conditioned on the prompt and its rubrics, enabling two complementary mechanisms: an Adaptive branch that reactively repairs reward misspecification; and a Chase branch that proactively converts f...
Si-Yuan Li, Xin-Xin Song, Rui-Nian Chen et al.· 0 citations
Experiments on multimodal clinical and cross-domain benchmarks demonstrate that EvtGraph outperforms both Transformer-based and recurrent baselines while significantly improving efficiency, suggesting that budget-constrained event-centric representation provides a general paradigm for learning from high-redundancy temp...
Zi-Qian Wang, Tingxiong Xiao, Yuxiao Cheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.