2607.21106v1 Jul 23, 2026 cs.AI

AttriMem: 설명력 기반 프로세스 피드백을 통한 에이전트 메모리 학습

AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction

Xuhong Zhang
Xuhong Zhang
Citations: 2
h-index: 1
Qinfeng Li
Qinfeng Li
Citations: 48
h-index: 3
Yuntai Bao
Yuntai Bao
Citations: 37
h-index: 2
Wenqiang Zhang
Wenqiang Zhang
Citations: 143
h-index: 2
Xinyang Yu
Xinyang Yu
Citations: 7
h-index: 1
Hongze Chen
Hongze Chen
Citations: 0
h-index: 0

효과적인 메모리는 LLM 에이전트에 매우 중요하지만, 이를 효과적으로 구축하는 것은 여전히 어려운 과제입니다. 메모리 구축 정책은 상호 작용이 축적됨에 따라 어떤 정보를 추출하고, 저장하고, 업데이트하고, 압축하거나 버릴지 결정합니다. 휴리스틱 기반 메모리 방법은 주관적인, 작업별 규칙에 의존하며, 이는 하위 목표와 일치하지 않을 수 있으며 작업 간의 적응성을 제한할 수 있습니다. 반면, 강화 학습(RL) 기반 방법은 작업 피드백을 통해 학습하지만, 주로 결과 수준 또는 모듈 수준의 보상을 사용합니다. 이러한 포괄적인 신호는 작업 성공을 나타내지만, 최종 답변을 지원하는 중간 메모리 내용을 식별하지 못하여 세분화된 책임 할당 병목 현상을 야기합니다. 그러나 그러한 프로세스 피드백을 구축하는 것은 매우 어렵습니다. 왜냐하면 중간 메모리 결정에는 고유한 정답 목표가 없으며, 적절한 기여도는 에이전트의 불확실한 추론 경로에 따라 달라지므로 사전에 지정할 수 없습니다. 본 논문에서는 강화 학습(RL)을 통해 메모리 구축 정책을 학습하기 위한 설명력 기반 프로세스 피드백 프레임워크인 AttriMem을 제안합니다. AttriMem은 전역 결과 보상에 더하여, 최종 답변에 대한 토큰 수준의 기여도를 기반으로 파생된 로컬 보상을 추가합니다. 장기 대화 질의 응답 실험에서 AttriMem은 검색 기반, 휴리스틱 기반 및 RL 기반 방법과 같은 기존 방법을 능가하며, 다양한 벤치마크 및 답변 모델을 통해 일반화되고, 강화 학습 최적화를 안정화시킵니다.

Original Abstract

Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with downstream objectives and limit cross-task adaptability. RL-based methods, by contrast, learn from task feedback but mainly use outcome- or module-level rewards. These coarse signals indicate task success but cannot identify which intermediate memory contents support the final answer, creating a fine-grained credit-assignment bottleneck. However, constructing such process feedback is prohibitively difficult because intermediate memory decisions lack unique ground-truth targets, while the appropriate credit varies with the agent's uncertain reasoning trajectory and therefore cannot be specified in advance. We propose AttriMem, an attribution-guided process-feedback framework for learning memory-construction policies with RL. AttriMem augments the global outcome reward with local rewards derived from token-level contributions to the final answer. Experiments on long-horizon dialogue question answering show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!