#artificial intelligence
May 2026
TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
TwinRouterBench is introduced, a step-level routing benchmark with two tracks that supports fast offline iteration followed by end-to-end validation under live agent execution, and success is measured by official task resolution and realized API spend.
Pei Yang, Wan-Yi Chen, Tong Yang et al.
· arXiv.org · 6 citations