Skip to content

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

Jul 2026 · arXiv.org · Vol abs/2607.23933 · 1 citation · 47 references
Computer Science

TL;DR

This work presents SpecBox, a runtime built around speculative sandbox preallocation tailored for dynamic LLM agent execution pipelines, and implements keyword matching and streaming semantic embedding to enable intent-driven sandbox prewarming, which identifies pending tool execution demands mid-LLM token generation and fully overlaps sandbox bootstrapping with model inference.

Abstract

As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency. Persistent long-lived sandbox reservations incur excessive memory overhead at scale, while lazy on-demand instantiation generates severe cold-start penalties that degrade response performance under multi-tenant, multi-turn agent workloads. To resolve this dilemma, we present SpecBox, a runtime built around speculative sandbox preallocation tailored for dynamic LLM agent execution pipelines. At its core, SpecBox implements keyword matching and streaming semantic embedding to enable intent-driven sandbox prewarming, which identifies pending tool execution demands mid-LLM token generation and fully overlaps sandbox bootstrapping with model inference. To extend prewarming windows across sequential agent steps, the framework leverages context-aware stochastic prefetching atop a sandbox dependency graph to probabilistically forecast future sandbox switches ahead of execution. We complement these speculative mechanisms with two orthogonal optimizations: a semantic result cache that prunes redundant repeated sandbox invocations, and a dedicated out-of-band shared-memory transport plane that bypasses conventional network serialization to deliver zero-copy artifact transfers. Evaluated on high-concurrency multi-turn agent traces, our prototype demonstrates that SpecBox cuts P99 end-to-end latency by up to $2.9\times$ relative to the on-demand sandbox baseline, while slashing peak memory consumption by $45.9\%$ compared to permanently reserved sandbox deployments.

View source

Similar papers

Preprint Aug 2026

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

PeakBench is a benchmark of executable multi-tool workflows with execution-grounded dependency annotations and measured resource profiles that shows that strong logical planning does not reliably translate into safe or efficient execution under resource constraints, and exposes resource information to reduce avoidable...

Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li et al. · 0 citations
Preprint Sep 2026

RR-Evict: Fine-Grained Prefix Cache Eviction beyond LRU for Agentic LLM Serving

LLM-based agents execute long-horizon tasks through repeated model calls interleaved with tool execution and user interaction. As each call extends the history accumulated in previous turns, prefix caching avoids repeated prefill of the agent's entire context. However, the aggregate cache footprint grows with context l...

Zai-Feng Pan, Chris Wu, Zheng-Ding Hu et al. · 0 citations
Preprint Sep 2026

PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale

Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing sched...

Zhi-Yuan Tan, De-Jiang Zhu, Jing-Zhe Jiang et al. · 0 citations
Preprint Aug 2026

TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving

A Task-Oriented Prefix-Aware Scheduler that jointly decides which agent prefixes to keep in the cache and which requests to schedule for execution and scores candidate post-decision states by trading off the expected reduction in each task's longest remaining service path against the near-term benefit of downstream pre...

Hongqiu Ni, Han Tian, Chi Zhang et al. · 2 citations
Preprint Aug 2026

Bridging Agent Semantics with Spot Capacity: An Elastic and Recoverable Service Model

LLM agents increasingly drive long-running cloud inference workloads in which model calls differ in urgency, redundancy, completion semantics, and replay cost. Model-as-a-Service (MaaS) platforms expose several service models for trading cost against latency, availability, and capacity commitment. These models operate...

Min-Chen Yu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.