2605.28069v1 May 27, 2026 cs.AI

ZipRL: 과거 응답 재활용을 통한 적응형 다중 턴 컨텍스트 압축

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

Li Wang
Li Wang
Citations: 8
h-index: 2
Guojun Yin
Guojun Yin
Citations: 265
h-index: 8
Xiaohan Wang
Xiaohan Wang
Citations: 101
h-index: 6
Jiajun Chai
Jiajun Chai
Citations: 116
h-index: 6
Zhexin Hu
Zhexin Hu
Citations: 0
h-index: 0
Xiaojun Guo
Xiaojun Guo
Citations: 32
h-index: 3
Wei Lin
Wei Lin
Citations: 57
h-index: 3

대규모 언어 모델(LLM)을 복잡하고 여러 단계로 구성된 에이전트 작업에 적용하기 위해서는 적응형 컨텍스트 압축이 필수적입니다. 그러나 규칙 기반 압축 방법은 작업 수행에 중요한 세부 사항을 삭제할 수 있으며, 강화 학습(RL) 접근 방식은 일반적으로 장기적인 워크플로우에서 발생하는 희소한 보상으로 인해 정보 유지와 토큰 효율성 간의 균형을 맞추는 데 어려움을 겪습니다. 이러한 격차를 해소하기 위해, 우리는 검증 가능한 보상을 활용한 강화 학습(RLVR)에 특화된 새로운 적응형 압축 프레임워크인 ZipRL을 제안합니다. ZipRL은 능동적이고 균일하지 않은 정보 감소를 위한 다중 수준의 압축 메커니즘과 함께, RLVR 최적화 과정에서 훈련 신호를 강화하도록 설계된 기술인 과거 응답 재활용(HRR)을 특징으로 합니다. 이론적으로 우리는 ZipRL이 균일한 방법에 비해 더 우수한 작업 관련 유틸리티를 제공한다는 것을 증명합니다. 구체적으로, ZipRL은 거시 압축을 위해 세분화된 프롬프트를 사용하고, 일반화된 장점 재형성을 통해 HRR을 GRPO에 통합합니다. 다양한 버전과 파라미터 규모의 여러 모델을 사용하여 제안하는 접근 방식의 효과를 검증했습니다. 다섯 가지 에이전트 작업에 대한 벤치마크 결과, ZipRL은 Qwen3-4B 및 Qwen3-8B 모델에서 최첨단 기술보다 각각 27.9%와 34.7% 더 높은 성능을 보였으며, 뛰어난 토큰 효율성과 극단적인 256턴 초과 추론 환경에서도 안정성을 유지했습니다.

Original Abstract

Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression methods may discard task-critical nuances, while Reinforcement Learning (RL) approaches usually struggle to balance information retention and token efficiency under the sparse rewards inherent to long-horizon workflows. To bridge this gap, we propose ZipRL, a novel adaptive compression framework tailored for Reinforcement Learning from Verifiable Rewards (RLVR). ZipRL features a multi-granularity compression mechanism for active, non-uniform information reduction, coupled with Hindsight Response Replay (HRR), a technique designed to densify training signals during RLVR optimization. Theoretically, we prove ZipRL's superior task-relevant utility over uniform methods. Concretely, ZipRL utilizes coarse-to-fine prompts for macro-compression and incorporates HRR into GRPO via generalized advantage reshaping. Multiple models of varying versions and parameter scales validate the effectiveness of our approach. Benchmarks on five agent tasks show ZipRL outperforms state-of-the-art approaches by 27.9% and 34.7% across Qwen3-4B and Qwen3-8B models, while maintaining exceptional token efficiency and robustness under extreme 256-turn extrapolation stress tests.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!