Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task re...
Jin-Hao Dong, Liang Zhao, Zi-Hao Yue et al.· 0 citations
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better...
Bo-Wen Ye, Lei Li, Shi-Cheng Li et al.· 0 citations
This work studies ViT attention heads and finds they differentiate into object- and background-specialist roles, a pattern most pronounced under full attention, and proposes SHS-Index to quantify this specialization, showing that it distinguishes full-attention from chunk-window ViTs, and finds that it strongly tracks...
Chenyu He, Lei Li, Shi-Cheng Li et al.· 0 citations
The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.
Zi-Yue Wang, Aomufei Yuan, Yi-Ran Yao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.