2607.24368v1 Jul 27, 2026 cs.CL

Keep It InMind: 에이전트 메모리에서 암묵적 연관성 인지 오류 현상에 대한 성능 측정

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Zhendong Mao
Zhendong Mao
Citations: 275
h-index: 7
Mingxuan Du
Mingxuan Du
Citations: 193
h-index: 3
Benfeng Xu
Benfeng Xu
Citations: 1,349
h-index: 15
Ruizhe Li
Ruizhe Li
School of Artificial Intelligence and Data Science, University of Science and Technology of China
Citations: 42
h-index: 3

장기 기억 시스템은 사용자가 말한 내용을 외부 저장소에 저장하고, 관련된 질문이 들어오면 이를 검색합니다. 이 인터페이스는 매우 당연하게 여겨져 명시적으로 언급되지 않는 가정을 기반으로 합니다. 즉, 필요한 기억은 해당 질문과 유사해야 한다는 것입니다. 하지만 일반 상식 지식은 이러한 가정을 깨뜨립니다. 예를 들어, 특정 음식에 대한 알레르기가 있는 사용자가 마카롱을 요청했을 때, 마카롱의 주재료인 아몬드 분 때문에 답변이 달라져야 하지만, 두 텍스트는 검색 시스템이 인지할 수 있는 단서가 전혀 없을 수 있습니다. 우리는 이러한 오류 현상을 '암묵적 연관성 인지 오류(implicit-association blind spot)'라고 부르며, 10가지 생활 영역을 포괄하는 125개의 작업으로 구성된 InMind라는 성능 측정 도구를 소개합니다. InMind은 전문가가 검증했으며, 113개 작업은 공개된 출처를 기반으로 합니다. InMind의 제어 그룹은 기존 평가에서 혼동되는 세 가지 설명을 분리합니다. 즉, 정보가 저장되지 않았거나, 모델이 연결 지식(bridging knowledge)이 부족하거나, 정보는 저장되었지만 검색되지 않은 경우입니다. 실험 결과, 관련 정보를 맥락에 맞게 제공하면 시스템은 84.0%의 간접적인 질문에 대해 정확하게 답변합니다. 반면, 동일한 정보를 검색해야 하는 경우, 벡터 기반, 그래프 기반 및 에이전트 기반 메모리 시스템은 최대 14.4%의 정확도를 보입니다. 이는 해당 정보가 필요할 때 즉시 100%까지 회상될 수 있음에도 불구하고 나타나는 현상입니다. 차원을 8배 늘린 임베딩을 사용하더라도 모든 시스템에서 목표 회상률이 향상되지만, 성능 격차는 거의 그대로 유지됩니다. 메모리가 질문 전에 계속 표시되도록 하는 간단한 진단 도구를 사용하면 대부분의 격차가 해결되며, 오류가 질문 처리 인터페이스 자체에 있다는 것을 확인하고, 어떤 정보를 항상 표시해야 할지 결정하는 라우팅(routing)이 중요한 문제임을 밝힙니다. InMind는 이러한 문제를 평가하기 위한 목적으로 설계되었습니다.

Original Abstract

Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!