Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended c...
Ji-Hua Tao, Xiao-Kun Yuan, Yao-Ming Li et al.· 0 citations
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better...
Bo-Wen Ye, Lei Li, Shi-Cheng Li et al.· 0 citations
A benchmark based on real-world MCP definitions designed to evaluate the tool-use capabilities of agents, which reveals significant performance differences in handling complex, multi-step tool invocations.
Zixiang Liu, Wenrui Liu, Elsie Dai et al.· arXiv.org· 10 citations
This work studies ViT attention heads and finds they differentiate into object- and background-specialist roles, a pattern most pronounced under full attention, and proposes SHS-Index to quantify this specialization, showing that it distinguishes full-attention from chunk-window ViTs, and finds that it strongly tracks...
Chenyu He, Lei Li, Shi-Cheng Li et al.· 0 citations
This work introduces PersonaForge, a user simulation framework for synthesizing realistic multi-turn user--agent interactions that combines a four-dimensional persona space, SOUL-driven behavioral control calibrated to real-user statistics, and Reverse Deep Construction grounded in authentic seed queries.
Hanglong Lv, Dawei Zhu, Lei Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.