Skip to content
Preprint

MidTool: Mid-training Data Synthesis for Agentic Tool Use

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

This work presents MidTool, an open corpus construction pipeline for agentic tool-use mid-training that combines large-scale web, PDF, and code data with synthesized supervision from real-world tool APIs, MCP skills, and document-grounded workflows and suggests that general tool use, like other important LLM capabilities, benefits from dedicated mid-training rather than being left entirely to post-training.

Abstract

Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Recent work has shown that targeted mid-training can strengthen reasoning-intensive abilities such as math and science, and can also improve agentic capabilities in software-engineering settings. In this work, we study the parallel but less explored agentic capability: general tool use. We present MidTool, an open corpus construction pipeline for agentic tool-use mid-training that combines large-scale web, PDF, and code data with synthesized supervision from real-world tool APIs, MCP skills, and document-grounded workflows. MidTool is designed to teach models how to recognize tool affordances, ground arguments from context, compose tool call workflow, and recover from incomplete information. We mid-train Qwen3-4B-Base and Qwen3-8B-Base on MidTool-Mix, and then apply follow-up post-training with both supervised fine-tuning and reinforcement learning. Compared with baselines, MidTool-Mix consistently improves downstream performance under both SFT and RL on BFCL, tau2-Bench, and MCP Universe. These results suggest that general tool use, like other important LLM capabilities, benefits from dedicated mid-training rather than being left entirely to post-training.

View source

Similar papers

Preprint Aug 2026

SPT: Skills as Pre-Training Data for Agentic Language Models

Agentic (tool-using) language models are mainly trained on tool-call traces and agent trajectories during post-training. These data provide direct behavioral supervision, but producing them requires task environments, execution, and verification, making broad tool and task coverage expensive. Publicly available skills...

Yufei Sun, Yudong Li, Yi-Min Cheng · 0 citations
#artificial intelligence Preprint Aug 2026

ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL

ReToolSQL is presented, a two-stage training framework for text-to-SQL that combines a supervised warm-start on rejection-sampled reasoning traces with agentic reinforcement fine-tuning (RFT) over multi-turn tool-use trajectories and shows that a properly designed SFT$\to-RFT pipeline over tool-use trajectories is a pr...

Pratik Kakkar, Chandra Dhir, Ravi Shankar et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development

Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic to...

Xin Yu, Li-Zhu Zhang, Jiamu Bai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee com...

Bo Mao, Hang He, Lin-Ting Wang et al. · 0 citations
Preprint Aug 2026

Joint Optimization of Tool Creation and Use for Large Language Model Agents

A reinforcement learning framework that jointly trains tool creation and tool use inside a single policy, with three separate reward axes that catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient.

Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen et al. · 2 citations
Book Open access Aug 2026

Recipes for Agents: Understanding Skills and Their Open Questions

This paper examines how skills may help address bottlenecks of current agents and how they may expand agent capabilities through reusable domain procedures loaded at inference time and outlines open questions in skill construction, composition, evaluation, portability, governance, and security.

Hanwen Xing, Haomin Zhuang, Xuandong Zhao et al. · 7 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.