2606.31650v1 Jun 30, 2026 cs.LG

ECHO: 에이전트 강화 학습에서 선택적 턴 메모리를 활용한 가지치기를 통한 행동 결정 및 추적을 통한 학습

ECHO: Prune to act, trace to learn with selective turn memory in agentic RL

Binbin Zheng
Binbin Zheng
Citations: 11
h-index: 2
Yuyang You
Yuyang You
Citations: 30
h-index: 2
Guanqun Zhao
Guanqun Zhao
Citations: 0
h-index: 0
Zijun Xie
Zijun Xie
Citations: 0
h-index: 0
Enlei Gong
Enlei Gong
Citations: 43
h-index: 2
Jihua Liu
Jihua Liu
Citations: 12
h-index: 2
Lingfeng Liu
Lingfeng Liu
Citations: 0
h-index: 0
Jiayao Tang
Jiayao Tang
Citations: 4
h-index: 1
Aoqi Hu
Aoqi Hu
Citations: 0
h-index: 0
Zeyu Chen
Zeyu Chen
Citations: 0
h-index: 0

장기적인 관점에서 작동하는 언어 기반 에이전트는 도구와 반복적으로 상호 작용하고, 제한된 컨텍스트 창 내에서 증거를 축적하며, 의사 결정을 내려야 합니다. 기존의 컨텍스트 관리 방법은 먼 과거 기록을 잘라내거나, 이전 대화 내용을 요약하거나, 압축된 메모리 상태를 선택함으로써 이러한 과정을 가능하게 합니다. 그러나 이러한 발전에도 불구하고 두 가지 주요 제한 사항이 존재합니다. 첫째, 대화 횟수가 증가함에 따라 역사적 관찰 내용이 점차 제거되거나 압축된 상태로 저장되어, 정책이 세밀한 증거를 재사용하기 어려워집니다. 둘째, 원래의 대화 내용을 더 이상 참조할 수 없게 되면 결과 기반 강화 학습은 성공적인 최종 답변을 뒷받침하는 증거와 정책 업데이트를 명시적으로 연결할 수 있는 경로를 잃게 됩니다. 이러한 문제를 해결하기 위해 우리는 ECHO라는 선택적 턴 메모리 프레임워크를 제안합니다. ECHO는 소스 인덱스를 활용한 재구성을 통해 역사 정보의 손실과 추적 가능한 학습을 동시에 개선합니다. 구체적으로, ECHO는 각 완료된 환경 단계를 압축하여 간결한 메모리 레코드로 저장하고, 이러한 레코드에서 선택적으로 정보를 추출하여 제한된 정책 컨텍스트를 구성하며, 선택된 소스 인덱스를 재사용하여 긍정적인 결과에 대한 보상을 성공적인 답변을 뒷받침하는 증거 및 선택 행동으로 연결합니다. 실험 결과, ECHO는 BrowseComp-Plus 데이터셋에서 43.4%의 정확도를 달성했으며, 이는 GRPO (28.9%) 및 요약 기반 방법인 SUPO (36.1%)보다 우수한 성능입니다. 또한, ECHO는 더 적은 횟수의 대화와 낮은 경로 볼륨을 사용합니다 (그림 1). 더욱이, 학습된 정책은 밀집형과 MoE 아키텍처 모두에서 멀티 객체 QA, 코드 생성 및 심층 정보 검색 벤치마크에 대한 제로샷 일반화 성능을 향상시킵니다.

Original Abstract

Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Existing context-management methods make such rollouts feasible by truncating distant history, folding past turns into summaries, or selecting compact memory states. However, these breakthroughs introduce two coupled limitations. First, as the number of turns grows, historical observations are progressively removed or collapsed into compressed states, making it harder for the policy to reuse fine-grained evidence. Second, once the original turns are no longer source-addressable, outcome-based RL loses an explicit path for aligning policy updates with the evidence that supported a successful final answer. To this end, we propose ECHO, a selective turn-memory framework that jointly addresses history collapse and traceable learning through source-indexed reconstruction. Specifically, ECHO compresses each completed environment turn into a compact memory record, reconstructs bounded policy contexts by selecting from these records, and reuses the selected source indices to route positive outcome credit to the evidence and selection actions that support successful answers. On BrowseComp-Plus, ECHO reaches 43.4% held-out accuracy, outperforming GRPO (28.9%) and the rolling-summary baseline SUPO (36.1%), while using fewer turns and lower trajectory volume than SUPO (Figure 1). Additionally, the trained policy improves zero-shot generalization across multi-objective QA, code generation, and deep information-seeking benchmarks on both dense and MoE backbones.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!