Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works ref...
Qi Yu, Rui-Zhong Qiu, Zhichen Zeng et al.· 0 citations
Two implications: the cheapest tool surface is jointly determined by task regime and agent design rather than by either axis alone, and the headline cost signal lives in cache-adjusted cost -- not pass rate, which is invariant across surfaces at the model sizes the authors evaluate.
Hong Yang, Qi Yu, Travis Desell· arXiv.org· 0 citations
This work proposes a principled VLM TTA method called \algname, and theoretically reveals that the InfoNCE loss can be neatly reformulated as a Wasserstein OT formulation, thereby unifying the objectives of the inference and adaptation of VLMs to achieve their mutual benefits.
Qi Yu, Zhichen Zeng, Katherine Tieu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.