Skip to content

Toolcompass: Guiding Tool Trialing, Not Suppressing It

Sep 2026 · 0 citations · 59 references
Computer Science

TL;DR

This work introduces ToolCompass, a post-training framework that guides tool trialing by organizing tool-call representations according to shared functions and jointly reduces intra-function variation across domains and increases inter-function separation.

Abstract

Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment. A key challenge is tool trialing, i.e., excessive trials waste the interaction budget, whereas selective trials enable exploration of unfamiliar tools. Existing outcome-based post-training leaves wasteful trials unguided, while turn-level supervision may suppress necessary exploration. We introduce ToolCompass, a post-training framework that guides tool trialing by organizing tool-call representations according to shared functions. Specifically, ToolCompass models each function class as a von Mises--Fisher distribution and jointly reduces intra-function variation across domains and increases inter-function separation. This structure transfers experience from seen tools to functionally similar unseen tools, directing exploration away from unrelated alternatives. ToolCompass requires no ground-truth call traces or unseen-tool access and incurs no inference overhead. Experiments on AppWorld and FTRL show consistent gains across GRPO, RFT, and DMPO. improves AppWorld OOD task success by up to 10.71 percentage points over vanilla post-training and performs best among competitive baselines on both benchmarks.

View source

Similar papers

Conference Sep 2026

Mitigating Tool Overuse for LLMs via Active Knowledge Boundary Probing

RADAR is proposed, an automated framework for knowledge boundary discovery and tool overuse mitigation that globally propagates labels and rewrites conflicts, yielding deployment-aligned decisions with fewer unnecessary tool calls.

Zhao-Yu Yang, Wenjun Ke, Yuan-Yao Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Toollery: Scaling LLM Agents to Thousands of Skills and Tools

As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for LLM selection. We present \textbf{Toollery...

Xiang-Xi Tian, Ran-Yun Guan · 0 citations
#artificial intelligence Preprint Sep 2026

CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins

Experiments with retrieval as the main tool on Qwen3-4B show that CoBRA improves tool-use efficiency and boundary-sensitive answer accuracy while maintaining strong performance on tool-dependent out-of-distribution questions.

Wen-Hao Zou, Xiang-Long Liu, Wendong Bi et al. · 0 citations
#natural language process... Preprint Sep 2026

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

This work proposes ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection, which introduces category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools.

Zhen-Long Dai, Xu-Jie Song, Zi-Tong Wang et al. · 0 citations
Preprint Aug 2026

Joint Optimization of Tool Creation and Use for Large Language Model Agents

A reinforcement learning framework that jointly trains tool creation and tool use inside a single policy, with three separate reward axes that catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient.

Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen et al. · 2 citations
#small language model Preprint Aug 2026

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

This work proposes StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility, and experiments show that explicit strategy utility differs substantially across task categories and dep...

Jing-Han Tan, Yuanzhe Wang, Lu Chen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.