Preprint
Jul 2026
SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference
SPORK (Self-sPeculative fORKing), a training-free controller that dispatches the speculated tool call early, overlapping its execution with the remaining chain-of-thought decode, and is orthogonal to token-level speculative decoding.
Huajun Bai, Weiwei Lv, Huichuan Zheng et al.
· 1 citation