Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task re...
Jin-Hao Dong, Liang Zhao, Zi-Hao Yue et al.· 0 citations
This work proposes BiVCoder, a diagnosis-driven multi-agent framework featuring a novel bidirectional code-test diagnosis mechanism, and introduces BiVCoder-SFT, a role-specific instruction fine-tuning scheme.
Xiaoyang Li, Jin-Hao Dong, Wenhang Shi et al.· Proceedings of the 32nd ACM...· 0 citations
Large Language Models (LLMs) have demonstrated remarkable potential in automated code generation. However, existing test-driven code generation and refinement frameworks are often hindered by the tests' quality: they typically treat self-generated tests as ground truth, leading to ineffective debugging loops where code...
Xiaoyang Li, Jinhao Dong, Wenhang Shi et al.· Proceedings of the 32nd ACM...· 0 citations
DBcover is proposed, an LLM-driven database test generation framework that performs white-box, code-aware SQL test generation through contextual reasoning, and substantially outperforms existing fuzzers.
Yan-Kai Rong, Shuang Liu, Jin-Hao Dong et al.· Proceedings of the 2026 IEEE...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.