TwinRouterBench is introduced, a step-level routing benchmark with two tracks that supports fast offline iteration followed by end-to-end validation under live agent execution, and success is measured by official task resolution and realized API spend.
Pei Yang, Wan-Yi Chen, Tong Yang et al.· arXiv.org· 6 citations
On-policy distillation (OPD) pays twice for each fresh batch: the student generates trajectories and a stronger teacher scores them. Existing methods improve which trajectories are scored and how the teacher signal is constructed, but usually consume it with one actor update. We introduce CLOOPD, a closed-loop framewor...
Keye Zheng, Han-Yu Li, Zhan Cheng et al.· 0 citations
BenchShield is presented, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation that grounds detection in a finite lifecycle model of an evaluation's reward-relevant events and achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.
Sheng-Han Zheng, Zong-Lin Di, Yimin Liu et al.· 0 citations
InfraBench is presented, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment and shows that even the strongest agent cannot secure a full score across all tasks.
Yuan Gao, Zeren Yang, Junnan Li et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.