This work introduces ToolCompass, a post-training framework that guides tool trialing by organizing tool-call representations according to shared functions and jointly reduces intra-function variation across domains and increases inter-function separation.
Abstract
Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment. A key challenge is tool trialing, i.e., excessive trials waste the interaction budget, whereas selective trials enable exploration of unfamiliar tools. Existing outcome-based post-training leaves wasteful trials unguided, while turn-level supervision may suppress necessary exploration. We introduce ToolCompass, a post-training framework that guides tool trialing by organizing tool-call representations according to shared functions. Specifically, ToolCompass models each function class as a von Mises--Fisher distribution and jointly reduces intra-function variation across domains and increases inter-function separation. This structure transfers experience from seen tools to functionally similar unseen tools, directing exploration away from unrelated alternatives. ToolCompass requires no ground-truth call traces or unseen-tool access and incurs no inference overhead. Experiments on AppWorld and FTRL show consistent gains across GRPO, RFT, and DMPO. improves AppWorld OOD task success by up to 10.71 percentage points over vanilla post-training and performs best among competitive baselines on both benchmarks.
RADAR is proposed, an automated framework for knowledge boundary discovery and tool overuse mitigation that globally propagates labels and rewrites conflicts, yielding deployment-aligned decisions with fewer unnecessary tool calls.
Zhao-Yu Yang, Wenjun Ke, Yuan-Yao Li et al.· Proceedings of the Thirty-Fi...· 0 citations
As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for LLM selection. We present \textbf{Toollery...
Experiments with retrieval as the main tool on Qwen3-4B show that CoBRA improves tool-use efficiency and boundary-sensitive answer accuracy while maintaining strong performance on tool-dependent out-of-distribution questions.
Wen-Hao Zou, Xiang-Long Liu, Wendong Bi et al.· 0 citations
This work proposes ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection, which introduces category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools.
Zhen-Long Dai, Xu-Jie Song, Zi-Tong Wang et al.· 0 citations
A reinforcement learning framework that jointly trains tool creation and tool use inside a single policy, with three separate reward axes that catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient.
Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen et al.· 2 citations
This work proposes StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility, and experiments show that explicit strategy utility differs substantially across task categories and dep...
Jing-Han Tan, Yuanzhe Wang, Lu Chen et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.