Skip to content
Preprint

DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents

Aug 2026 · 0 citations · 56 references
Computer Science

TL;DR

Results show that DocsChisel improves the task success rate of LLM agents by 95.89% over the original tool documentation and by 75.15%, on average, over existing baselines, while incurring limited optimization time and token overhead.

Abstract

Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the tool-use capabilities of LLM agents, while largely treating tool documentation as a fixed input. Although several recent works attempt to optimize tool documentation through rewriting or compression, little is known about how the information contained in tool documentation affects agent performance across different settings. To bridge this gap, we conduct a large-scale empirical study on tool documentation for LLM agents. Our study reveals substantial heterogeneity in the information fields provided by existing tool documentation. Moreover, the effectiveness of different information fields is highly dependent on the task domain, LLM backbone, and agent paradigm, indicating that no fixed tool documentation can consistently generalize across diverse agent settings. Motivated by these findings, we propose DocsChisel, an adaptive tool documentation optimization framework for LLM agents. DocsChisel analyzes failed execution traces of a target LLM agent to identify documentation-related issues, and iteratively optimizes tool documentation by adding, removing, and refining information fields for each tool. We evaluate DocsChisel against two state-of-the-art baselines, i.e., EasyTool and DRAFT. Experimental results show that DocsChisel improves the task success rate of LLM agents by 95.89% over the original tool documentation and by 75.15%, on average, over existing baselines, while incurring limited optimization time and token overhead

View source

Similar papers

Preprint Aug 2026

ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents

Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to invoke external systems and complete tasks beyond text generation. However, clean end-to-end (E2E) success cannot identify where a tool-use failure originates or how it propagates through a call. We introduce ToolRobustBench, a stage-wise diagnostic benchmark for tool-calling agents, where a tool-calling agent is an LLM system that selects a tool, supplies structured arguments, and interprets its returned feedback. ToolRobustBench aligns four perturbation families with the tool-use pipeline: tool-interface, user-intent, tool-output/observation, and runtime-environment perturbations. It attributes failures to tool selection, schema grounding, argument binding, tool-output/runtime-feedback handling, and E2E task success. Experiments on 15,456 single-family instances across 7 models, 16 sampled local tools, 4 perturbation families, and 14 subtypes show high but non-uniform clean performance and substantial robustness degradation, with tool-output/observation perturbation the dominant bottleneck. Mixed-family experiments reveal non-additive failure patterns that are not explained by isolated single-family results. Thus, ToolRobustBench provides a deterministic and cascade-aware benchmark for diagnosing robustness beyond clean tool-calling accuracy;

YiShan Zheng, Yuan Wu, Yi Chang · 0 citations
Preprint Jul 2026

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

E-Bench is introduced, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting, and it shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code extension, reliability remains below 70%.

Weihuang Zheng, Tianyuan Zou, Eileen Ye et al. · 1 citation
Preprint Jul 2026

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents

SyntheticAgentTraceQA is proposed, an execution- first framework for generating scalable supervision data for tool- augmented agents and shows that execution-grounded supervision improves tool execution behavior, reference-trace agreement, and answer-generation performance on the evaluated tasks.

Hafsa Ouajdi, Francesco Giannuzzo, Alaa Boukhary et al. · 1 citation · ⚡1
Preprint Jul 2026

ToolAtlas: Learning Once, Reusing Everywhere with Tool-Side Memory

This work introduces ToolAtlas, a graph-based framework that builds a persistent provider-side tool memory of tool capabilities, failure boundaries, and cross-tool compositions through execution-verified probing and establishes provider-side tool memory as an effective and reusable paradigm for tool servers.

Yue Fang, Zhibang Yang, Fangkai Yang et al. · 0 citations
Book Open access Jul 2026

TestAgent: A Multi-Agent LLM Framework for Repository-Level Unit Test Generation

TestAgent, a multi-agent tool implemented as a VS Code extension that automates the generation of high-quality unit tests for Java projects using repository-level Code Knowledge Graphs, demonstrates its practical utility for regression testing and bug discovery.

Ye Shang, Quanjun Zhang, Zheng Zhan et al. · 0 citations
Open access Aug 2026

A tool selection mechanism for LLM-based model management agents

In Model-Based Engineering (MBE), practitioners frequently have to choose appropriate tools from many different options. LLM-based agents are software components that depend on Large Language Models (LLMs) to autonomously select and apply software tools to perform specific tasks. Although LLMs have already been used in the MBE context, LLM-based agents to assist users of MBE tools remain underexplored. This is particularly true in industrial environments where only medium-sized on-premise LLMs can be considered due to policies related to security or data privacy. To investigate the potential of LLM-based agents for MBE, we study model management tools for transformation and editing. Currently, off-the-shelf agents such as Microsoft Copilot can invoke model management tools when the task is explicitly described. However, these agents struggle to select the correct transformation or operation when they only have limited contextual information, especially when coupled with medium-sized LLMs. To overcome this, we propose an approach based on complementary mechanisms. First, we provide a server and associated LLM-based agent with dedicated tools for each transformation available on this server. We also provide a similar server and agent for model editing operations. Then, to enable the two agents to efficiently select transformations and operations (respectively), we rely on a tool retrieval technique based on a tool relevance score computed by an LLM. We evaluate these agents using generated model management datasets that we contribute to the community. The obtained results show that our LLM-based agents select more accurately modeling tools to perform user instructions than off-the-shelf agents.

Zakaria Hachm, Théo Le Calvar, Hugo Bruneliere et al. · 0 citations