SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving
This work presents SpecBox, a runtime built around speculative sandbox preallocation tailored for dynamic LLM agent execution pipelines, and implements keyword matching and streaming semantic embedding to enable intent-driven sandbox prewarming, which identifies pending tool execution demands mid-LLM token generation a...