EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li, Chaofan Tao, Yanli Wang, Huanxin Lin, Kwang-Ting Cheng, Chi Ying Tsui, Haoli Bai

arXiv 2026arXiv preprintFirst author

EfficientAgent overview

LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the same coding-agent workload it speeds up one deployment, slows down another, and changes nothing on a third. The reason is that cached state must survive until it is used again: an agent’s prefix is reused only if the host tier holds the reusable context of the whole agent pool, which we call the reuse working set.

EfficientAgent sizes and manages the host tier by this working set. A stack-distance model estimates the working set from agent histories to size the host tier, and a runtime policy stops writing large refills of evicted context when the tier is too small. On SWE-bench Verified coding agents, a host tier sized to the estimated working set cuts recomputed prompt tokens by 93% and end-to-end time by 39%.

Citation

Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li, Chaofan Tao, Yanli Wang, Huanxin Lin, Kwang-Ting Cheng, Chi Ying Tsui, and Haoli Bai. EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents? arXiv preprint arXiv:2609.33762, 2026.