Web agents are usually evaluated in live environments, where environment state and judge models drift between runs, so the same checkpoint rarely reproduces the same score, making controlled studies of training phenomena impractical. We present WebMRE, an offline benchmark of 541 tasks and 5,293 steps derived from succ...
Cheng-Guang Gan, Yun-Hao Liang, Qing-Hao Zhang et al.· 0 citations
A structured scoping survey organized around the question of what decision a test changes is presented, and distinguishes the Red--Green--Refactor cycle from test-conditioned generation, execution-guided refinement, test-mediated analysis, and evaluation-only testing.
Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al.· 0 citations
Generating tests from a natural-language specification requires both an input that exposes faulty behavior and a correct expected output. These requirements need not improve together: a model can increase test correctness by choosing easier inputs, or discover useful inputs whose expected outputs it cannot predict. We...
Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al.· 0 citations
This work studies test utilization through matched prompting controls, paired semantic interventions, and test suites selected by fault detection to make test-specified rule changes measurable alongside implementation capability and benchmark correctness.
Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al.· 1 citation
Instruction-tuning data are judged by quality metrics, and tuned models are judged by benchmarks, but both judgments pass through an output interface: the surface format in which an answer is written. Using gradient signatures across 12 tasks, four semantically equivalent interfaces, three model families, and controlle...
Cheng-Guang Gan, Han-Jun Wei, Yun-Hao Liang et al.· 0 citations
The Mutual Reinforcement Effect is tested in multimodal document understanding on three corpora, two of receipts and one of scanned business forms, comparing single-task, joint and conditioned training, which puts one granularity's gold output in the other's prompt during training only.
Cheng-Guang Gan, Yun-Hao Liang, Han-Jun Wei et al.· 0 citations
This work asks whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether it mostly reshapes behavior the supervised model already has, and explains the failure of GRPO.
Public tests are widely used to guide large language model code generation, but whether models treat them as executable specifications or merely as extra prompt context remains unclear. We study test-driven code generation on HumanEval+, MBPP+, and recent LiveCodeBench tasks using Qwen2.5-Coder-7B and Qwen3.6-27B. We c...
Yunhao Liang, Chengguang Gan, Ruixuan Ying et al.· 0 citations
MAG is introduced, the first benchmark that unifies task execution and guide writing into a single Multimodal Action and Guide task, with two grounding schemes over screenshots: Set-of-Mark element selection and raw pixel coordinates.
The results show that executable feedback can repair secure-code generation, but its benefits depend on the model, task, feedback entry point, and especially test coverage.
Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al.· 1 citation
An audit-and-placebo protocol is proposed that separates verifier artifacts, interaction scaffolding, and grounded feedback credit in evaluations of self-evolving test generators in evaluations of self-evolving test generators.
Yunhao Liang, Chengguang Gan, Ruixuan Ying et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.