2606.17016v1 Jun 15, 2026 cs.CL

TokenPilot: LLM 에이전트를 위한 효율적인 캐시 기반 컨텍스트 관리

TokenPilot: Cache-Efficient Context Management for LLM Agents

Xuehai Wang
Xuehai Wang
Citations: 7
h-index: 2
Ning Zhang
Ning Zhang
Citations: 239
h-index: 6
Yunzhi Yao
Yunzhi Yao
Zhejiang University;Shandong University
Citations: 3,270
h-index: 22
Xinle Deng
Xinle Deng
Citations: 141
h-index: 6
Buqiang Xu
Buqiang Xu
Citations: 15
h-index: 2
Jizhan Fang
Jizhan Fang
Citations: 124
h-index: 4
Chiyu Wu
Chiyu Wu
Citations: 3,270
h-index: 4
Z. Xue
Z. Xue
Citations: 17
h-index: 2
Dian Chen
Dian Chen
Citations: 25
h-index: 4
C. Fu
C. Fu
Citations: 9
h-index: 1
Caiying Huang
Caiying Huang
Citations: 0
h-index: 0
Chenyu Jiang
Chenyu Jiang
Citations: 10
h-index: 1
Yijun Chen
Yijun Chen
Citations: 37
h-index: 2
Jingbo Shang
Jingbo Shang
Citations: 32
h-index: 2
Gong Yu
Gong Yu
Citations: 0
h-index: 0

LLM 에이전트가 장기 세션에서 운영될 때, 컨텍스트 축적은 추론 비용을 증가시키는 요인이 됩니다. 기존 방법들은 텍스트 가지치기 또는 동적 메모리 제거를 사용하여 토큰 사용량을 줄이지만, 제약 없는 시퀀스 변경은 레이아웃을 변경하여 접두사 불일치를 발생시키고 캐시 유효성을 손상시킵니다. 이는 텍스트 희소성과 프롬프트 캐시 연속성 간의 중요한 상충 관계를 드러냅니다. 이러한 문제를 해결하기 위해, 우리는 이중 수준의 컨텍스트 관리 프레임워크인 TokenPilot을 제안합니다. 전역적으로, Ingestion-Aware Compaction은 프롬프트 접두사를 안정화하고 입력 단계에서 외부 환경 노이즈를 제거하는 역할을 합니다. 지역적으로, Lifecycle-Aware Eviction은 컨텍스트 세그먼트의 잔여 유용성을 지속적으로 모니터링하여 작업 관련성이 만료될 때만 콘텐츠 세그먼트를 안전하게 제거하는 배치 업데이트 스케줄을 적용합니다. PinchBench 및 Claw-Eval 데이터셋에 대한 실험 결과, TokenPilot은 독립 실행 모드에서 61%와 56%의 비용 절감 효과를 보였으며, 연속 실행 모드에서는 각각 61%와 87%의 비용 절감 효과를 보였습니다. 성능은 기존 시스템과 비교하여 경쟁력 있는 수준을 유지했습니다. TokenPilot은 LightMem2 (https://github.com/zjunlp/LightMem2)에 통합되었습니다.

Original Abstract

As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approaches utilize text pruning or dynamic memory eviction to minimize token footprints; however, their unconstrained sequence mutations alter layouts, introducing prefix mismatches and cache invalidation. This reveals a critical trade-off between text sparsity and prompt cache continuity. To address this, we present TokenPilot, a dual-granularity context management framework. Globally, Ingestion-Aware Compaction acts as a framework harness to stabilize prompt prefixes and eliminate open-world environmental noise at the ingestion gate. Locally, Lifecycle-Aware Eviction monitors the ongoing residual utility of context segments, enforcing a conservative batch-turn schedule to offload content segments only when task relevance expires. Experiments on PinchBench and Claw-Eval under both isolated and continuous modes demonstrate that TokenPilot reduces costs by 61% and 56% in isolated mode, and 61% and 87% in continuous mode, while maintaining competitive performance compared to prior systems. TokenPilot has been integrated into LightMem2 at https://github.com/zjunlp/LightMem2.

1 Citations
0 Influential
51.036665926162 Altmetric
6.9 Score
Original PDF
54

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!