We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq endpoints and a thinking-enabled Gemini setting: ever-flip fractions are 0.7%, 2.0%, and...
Yi-Hang Chen, Pinyan Qian, Su Wang et al.· 0 citations
LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task d...
Yu Cheng, Yong-Kang Hu, Shuai-Jie Ma et al.· 0 citations
This work proposes Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving (TMCS), a step-by-step multi-agent framework that formalizes chemical problem solving as an interpretable, tool-augmented workflow.
Sheng-Qin Wang, Jie Jin, Yu Cheng et al.· 0 citations
Budget-aware evaluation for Active RAG evaluation is studied by recasting active retrieval as utility estimation, where retrieval is valuable only through its marginal correctness change over a no-retrieval answer.
Pinyan Qian, Su Wang, Chong Peng et al.· arXiv.org· 3 citations· ⚡1
BAP-SQL is presented, which treats observation formation as a budget-control stage: it estimates query risk, rewrites SQL when useful, and delegates hard limits to an independent runtime shield and improves tight-budget success.
Chong Peng, Pinyan Qian, Su Wang et al.· 2 citations
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator-worker inte...
Yi-Hang Chen, Yu-Xiang Chen, Yuxuan Huang et al.· 0 citations
This work introduces the Counterfactual Fabrication Lab, a deterministic micro-lab where the correct action is known: do nothing, and presents the Counterfactual Fabrication Lab for measuring fabricated failures in self-improving agent harnesses.
Su Wang, Pinyan Qian, Yifan Lin et al.· arXiv.org· 6 citations· ⚡2
This work introduces the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection.
Zhong-Wei Yu, Yan Song, Xue Yan et al.· 0 citations
Family-level and seed-stability analyses show that example-level accuracy overstates counterfactual consistency: complete four-way family success is rare, and an exploratory follow-up that elicits decomposed semantic evidence fails to improve routing for the cleanly evaluated endpoint.
Yi-Hang Chen, Pinyan Qian, Su Wang et al.· 0 citations
A minimal benchmark design and candidate reporting metrics for user-conditioned adaptation are proposed and a concrete design requirement for future personal-agent evaluation, with metrics used as reporting tools for that requirement.
Pinyan Qian, Su Wang, Yihang Chen et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.