Overall, corpus growth increasingly favors global candidate ranking: lexical retrieval is the strongest scalable default, while agentic reasoning works best after ranked discovery rather than in place of it.
Pengyu Wang, Benfeng Xu, Shaohan Wang et al.· arXiv.org· 1 citation
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement...
Yuhan Zhu, Ming-Xuan Du, Benfeng Xu et al.· arXiv.org· 0 citations
The results indicate that for precise, evidence-grounded questions over chat archives, much of the benefit credited to elaborate memory structures is recoverable by giving an agent controllable search over the unmodified record, with no LLM-based index construction at all.
Ruizhe Li, L. Zhang, Benfeng Xu et al.· 1 citation
GraphSynthQA, a knowledge-graph)—guided synthesis framework in an open-web setting, which iteratively retrieves and verifies evidence from the internet to expand a KG, then synthesizes complex, answer-verifiable queries grounded in multi-evidence dependencies.
Chiwei Zhu, Mingxuan Du, Benfeng Xu et al.· Annual International ACM SIG...· 0 citations
This work introduces InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources, and calls this failure mode the implicit-association blind spot, and introduces a minimal diagnostic probe that keeps memory visible before the query arrives recovers most of...
Ruizhe Li, Mingxuan Du, Benfeng Xu et al.· arXiv.org· 0 citations
This work introduces MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities, and proposes MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction.
Jiajia Lin, Mingxuan Du, Tuowen Zhou et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.