SubtleMemory: 장기적인 인공지능 에이전트의 미세한 관계적 기억 구분을 위한 벤치마크
SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents
OpenClaw와 같은 지속적인 인공지능 어시스턴트는 장기간 상호작용을 통해 방대한 양의 관련된 기억들을 축적합니다. 이러한 기억들이 증가함에 따라, 서로 강화되거나, 맥락에 따라 분기될 수 있으며, 심지어 직접적으로 충돌할 수도 있습니다. 따라서 정확한 지원은 개별적인 회상 능력보다는 기억 간의 관계를 이해하는 데 달려있습니다. 기존의 장기 기억 벤치마크는 이러한 관계가 다운스트림 작업에서 어떻게 유지되고 활용되는지를 탐구하기 어렵습니다. 이러한 격차를 해소하기 위해, 우리는 장기간 실행되는 인공지능 에이전트의 미세한 관계적 기억 구분을 위한 벤치마크인 SubtleMemory를 소개합니다. SubtleMemory는 관계 제어된 잠재 의미적 요소들을 구성하며, 이 요소들의 변형은 상호 보완적이거나, 미묘하게 다르거나, 모순적인 관계를 나타냅니다. 이러한 요소들은 현실적인 사용자-에이전트 대화 기록에 내장되어 있으며, 에이전트는 이후 쿼리와 지시 사항을 처리하는 동안 분산된 관계 구조를 복구해야 합니다. 이 벤치마크는 10개의 장기 기록에 걸쳐 1,522개의 평가 인스턴스를 포함하며, 1,090개의 관계 제어된 기억 변형 집합을 기반으로 하며, 사용자 관련 및 비사용자 관련 쿼리를 모두 다룹니다. 독립적인 6개의 메모리 시스템, 네이티브 메모리 모듈을 사용하는 두 개의 Claw 스타일 에이전트, 그리고 플러그인 메모리 모듈을 사용하는 세 개의 Claw 스타일 에이전트를 평가한 결과, 현재 시스템은 미세한 관계적 기억 구별 능력에서 여전히 취약하다는 것을 확인했습니다. 또한, 우리는 다양한 메모리 보존, 검색 및 다운스트림 추론 단계에서의 기능 프로필 차이를 보여주는 진단 프로토콜을 추가로 제시합니다.
Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they may reinforce one another, diverge across contexts, or directly conflict, making correct assistance depend on memory relations rather than isolated recall. Existing long-term memory benchmarks rarely probe how agents preserve and utilize such relations during downstream tasks. To address this gap, we introduce SubtleMemory, a benchmark for fine-grained relational memory discrimination in long-running AI agents. SubtleMemory constructs relation-controlled latent semantic artifacts whose variants instantiate complementary, nuanced, or contradictory relations, and embeds them into realistic user-agent histories, requiring agents to recover distributed relational structures during later queries and instructions. The benchmark contains 1,522 evaluation instances over 10 long histories, grounded in 1,090 relation-controlled memory-variant sets and spanning user-related and non-user-related queries. Evaluating six standalone memory systems, two Claw-style agents with native memory modules, and three Claw-style agents with plugin memory modules, we find that current systems remain weak on fine-grained relational memory discrimination. We further introduce diagnostic protocols that reveal distinct capability profiles across memory preservation, retrieval, and downstream reasoning stages.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.