Long-running tools can dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent idles. This observation stall presents the same tension that drove out-of-order processors -- asequential interface hides work that can be predicted and started early, but a spec...
Jiang-Nan Yu, Ce-Yu Xu, Meng-Ming Li et al.· 0 citations
LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the...
Kun-Ming Shao, Jie-Run Chen, Jiang-Nan Yu et al.· 0 citations
AgentZip is presented, the first memory compression system designed specifically for AI-agent sandboxes, which broadens the compression scope to any page with a profitable representation and shifts overhead control from compression-time page selection to restore-time prefetching.
Meng-Ming Li, Ce-Yu Xu, Qi-Jun Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.