Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended c...
Ji-Hua Tao, Xiao-Kun Yuan, Yao-Ming Li et al.· 0 citations
BackendForge is introduced, a benchmark of 56 contract-defined backend generation tasks rewritten from real open-source applications that suggests that current LLMs can implement many local API behaviors, but still struggle to produce complete backend services.
Yuzhe Guo, Mengzhou Wu, Yuan Cao et al.· arXiv.org· 0 citations
This work analyzes real chatbot failures to identify six recurring mechanisms and defines six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding, which shows that Hy-MultiTurn is broadly challenging.
Eileen Ye, Ji-Hua Tao, Yao-Ming Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.