LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like cl...
Qiong-Qiong Cao, Kang-Ni Liu, Xuan Kan et al.· 0 citations
Autonomous agents that automatically build artificial intelligence (AI) models could broaden access to AI across science and engineering. A popular line of such agents frames model building as a code search problem and solves it by tree search, in which each node is a candidate program and the tree grows by generating...
Pei-Jia Qin, Rui-Yi Zhang, Qi Cao et al.· 0 citations
Retrieval assembles repository context by ranking passages for relevance to the current query. A coding agent halfway through an issue has already read much of what such a ranker returns. Relevance is scored per passage, but sufficiency belongs to the set: independently scored passages can fill the budget with support...
Zhe-Xi Feng, Rui-Yi Zhang, Yong-Bo Yang et al.· 0 citations
Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings understate a harder assistant-memory regime: a fla...
Zhe-Xi Feng, Rui-Yi Zhang, Yong-Bo Yang et al.· 0 citations
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation ph...
Vivek Chavan, Peng-Tao Xie, Ya-Huan Shi et al.· 0 citations
Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators.
Xin-Yue Zhao, Rui-Yi Zhang, Li-Qin Ye et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.