2607.28272v1 Jul 30, 2026 cs.AI

MemHarness: 기억은 재생되는 것이 아니라 재구성된다

MemHarness: Memory Is Reconstructed, Not Replayed

Daocheng Fu
Daocheng Fu
Citations: 1,691
h-index: 20
Jianbiao Mei
Jianbiao Mei
Citations: 915
h-index: 16
Rong Wu
Rong Wu
Zhejiang University
Citations: 128
h-index: 5
Xuemeng Yang
Xuemeng Yang
Citations: 507
h-index: 12
Pinlong Cai
Pinlong Cai
Citations: 2,009
h-index: 19
Licheng Wen
Licheng Wen
Citations: 1,214
h-index: 15
Botian Shi
Botian Shi
Citations: 561
h-index: 13
Yu Yang
Yu Yang
Citations: 73
h-index: 3
Tao Hu
Tao Hu
Citations: 43
h-index: 3
Shu Zou
Shu Zou
Citations: 0
h-index: 0
Hairong Zhang
Hairong Zhang
Citations: 9
h-index: 2
Yuxin Wang
Yuxin Wang
Citations: 0
h-index: 0
Cong Zhang
Cong Zhang
Citations: 124
h-index: 6

과거 경험을 활용하는 것은 대규모 언어 모델 에이전트의 성능을 향상시키는 일반적인 전략으로 자리 잡았습니다. 그러나 대부분의 기존 메모리 기반 에이전트는 검색된 과거 경험을 정적 기록으로 취급하여, 에이전트의 현재 상황에 부합하지 않을 수 있음에도 불구하고 그대로 재생(replay)합니다. 이러한 '재생' 방식은 저장된 경험의 추상적이고 일반적인 특성과 의사 결정 시점에서 발생하는 구체적이고 끊임없이 변화하는 상태 사이의 간극을 간과하여, 종종 부정적인 영향을 초래합니다. 반면, 인간은 과거 경험을 그대로 회상하기보다는, 검색된 기억을 재구성하고 현재 맥락에 맞게 적용합니다. 이러한 점에 착안하여, 우리는 LLM 에이전트가 현재 맥락에 따라 과거 경험을 능동적으로 활용하고 재구성할 수 있도록 하는 프레임워크인 MemHarness를 제안합니다. 각 의사 결정 단계에서, 통합된 정책 모델은 현재 상태에 조건부로 검색된 경험을 비판적으로 검토하고 재구성하여, 행동하기 전에 맥락 기반의 지침을 제공합니다. 이러한 재구성 능력은 GRPO를 사용한 엔드투엔드 학습을 통해 자연스럽게 나타납니다. ALFWorld 및 WebShop에서의 실험 결과, MemHarness는 순수 강화학습(RL)과 정적 메모리 기반의 기존 방법보다 훨씬 뛰어난 성능을 보이며, 특히 데이터 분포가 다른 환경(OOD)에서 높은 안정성을 보여줍니다. 또한, 분석 결과에 따르면, 이러한 재구성 목표는 부정적인 영향을 방지할 뿐만 아니라 학습 과정 동안 잠재적인 지침 역할을 수행하여 에이전트의 근본적인 추론 능력을 향상시킵니다.

Original Abstract

Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent's current situation. This ``replay'' paradigm ignores the gap between the abstract, general nature of stored experience and the concrete, ever-changing states encountered at decision time, frequently causing negative transfer. In contrast, humans rarely recall past experiences verbatim; instead, they reorganize and adapt retrieved memories to fit the present context. Inspired by this, we propose MemHarness, a framework that equips LLM agents to actively harness and reconstruct past experiences based on the present context. At each decision step, a unified policy model critiques and reconstructs the retrieved experience conditioned on the current state, producing context-grounded guidance before acting. This reconstructive ability emerges naturally through end-to-end training with GRPO. Experiments on ALFWorld and WebShop show that MemHarness substantially outperforms pure RL and static memory-augmented baselines, demonstrating strong robustness in out-of-distribution (OOD) scenarios. Furthermore, our analyses reveal that this reconstruction objective not only prevents negative transfer but also serves as latent guidance during training, fundamentally improving the agent's intrinsic reasoning capabilities.

1 Citations
0 Influential
10 Altmetric
51.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!