수동적인 검색에서 능동적인 기억 탐색으로: 기억을 구조화된 행동 공간으로 활용하는 학습
From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space
개인화된 대화형 에이전트에게 장기 사용자 기억은 필수적이지만, 많은 기억 시스템들이 여전히 수동적인 검색 인터페이스를 통해 기억을 제공하여 모델이 미리 선택된 정보를 소비하게 만듭니다. 본 논문에서는 NapMem이라는 프레임워크를 소개합니다. 이는 장기 사용자 기억을 단순히 회수되는 맥락이 아닌 구조화된 행동 공간으로 활용하는 방법을 학습하는 것입니다. NapMem은 사용 기록을 연결된 다중 계층 메모리 피라미드로 구성하며, 원본 대화, 입력된 메모리 레코드, 주제 추적 및 사용자 프로필을 출처 관계를 통해 연결하고, 이러한 수준들을 메모리 도구를 통해 제공합니다. 에이전트는 쿼리와 중간 증거에 따라 기억을 선택하도록 학습되며, 답변하기 전에 다양한 메모리 세부 수준을 검토할 수 있습니다. PersonaMem-v2, LongMemEval 및 LoCoMo에서의 실험 결과는 NapMem 에이전트가 메모리 기반 강화 학습으로 훈련되었을 때 다양한 메모리 집약적인 작업에서 경쟁력 있는 성능을 보인다는 것을 보여줍니다. 또한, 메모리를 사용하지 않는 작업에 대한 평가 결과는 학습된 정책이 일반적인 추론 및 도구 사용 능력을 대체로 유지한다는 것을 시사합니다. 추가 분석에서는 저장 공간, 추론 비용, 도구 사용 행동 및 탐색, 메모리 세분화 및 강화 학습 훈련에 대한 요소 제거 실험을 검토합니다. 이러한 결과는 장기 사용자 기억이 적절한 세부 수준에서 기억을 활용하기 위한 학습된 정책과 구조화된 저장 방식을 결합함으로써 이점을 얻을 수 있음을 시사합니다.
Long-term user memory is essential for personalized conversational agents, yet many memory systems still expose memory through passive retrieval interfaces, making the model a consumer of pre-selected evidence. We introduce NapMem, a framework for learning to use long-term user memory as a structured action space rather than passively retrieved context. NapMem organizes user history into a linked multi-granularity memory pyramid, where raw conversations, typed memory records, topic tracks, and user profiles are connected through provenance relations, and exposes these levels through memory tools. The agent is trained to select memory according to the query and intermediate evidence, allowing it to inspect different memory granularities before answering. Experiments on PersonaMem-v2, LongMemEval, and LoCoMo show that a NapMem agent trained with memory-tool reinforcement learning is competitive across diverse memory-intensive tasks, while evaluations on non-memory tasks suggest that the learned policy largely preserves general reasoning and tool-use abilities. Additional analyses examine storage, inference cost, tool-use behavior, and ablations over navigation, memory granularity, and RL training. Our results suggest that long-term user memory benefits from coupling structured storage with a learned policy for using memory at the appropriate granularity.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.