기억이 권위가 될 때: 기억 통합 경계에서의 권위 붕괴 측정
When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary
지속적인 메모리는 (자기 진화하는) LLM 에이전트가 다양한 상호작용 기록을 재사용 가능한 사실, 선호도, 관찰 및 규칙으로 통합함으로써 여러 작업을 통해 적응할 수 있도록 합니다. 그러나 통합은 또한 암묵적인 권한 부여 경계를 설정하며, 저장된 정보가 나중에 사용자 사실, 증명된 관찰 또는 고정 지침으로 사용될 수 있는지 여부를 결정합니다. 본 연구에서는 권위 붕괴 현상을 밝히는데, 이는 통합 과정에서 주장이 유지되는 동시에 해당 주장의 합법적인 사용을 규제하는 출처 제약 조건이 삭제되어, 저장된 메모리가 실제 출처가 허용하는 것보다 더 큰 권한을 갖게 되는 현상입니다. 우리는 AuthMem-Bench라는 통제된 쌍대 벤치마크를 소개합니다. 이 벤치마크는 핵심 주장을 고정하고 다운스트림 작업도 고정한 상태에서, 오직 출처 권한만 변경하여 작성 시 붕괴, 다운스트림 권한 오류 및 자동 권한 보존을 평가합니다. 널리 사용되는 에이전트-메모리 시스템을 기반으로 한 일곱 가지 통합 모델과 LLM 백본을 사용하여 49개의 구성 중 48개에서 권위 붕괴 현상을 관찰했습니다. 통제된 동작 기반 평가에서는 권한 메타데이터가 없는 붕괴된 메모리의 평균 무단 동작 비율이 50.3%였습니다. 엔드 투 엔드 평가에서는 자동으로 예측되고 유지되는 권한 레이블이 관찰된 무단 동작 비율을 16.9%에서 0.0%로 줄였으며, 유해한 작업 성공률은 거의 변하지 않았습니다. 이러한 결과는 메모리를 기반으로 한 적응 과정에서 학습된 내용뿐만 아니라 해당 내용을 재사용할 수 있는 권한도 유지해야 함을 보여줍니다.
Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusable facts, preferences, observations, and rules. Yet consolidation also imposes an implicit authorization boundary: it determines whether stored information may later be consumed as a user fact, an attested observation, or a standing instruction. We identify authority collapse, in which consolidation preserves a claim while erasing the source constraints governing its authorized use, causing the stored memory to imply greater authority than its source permits. We introduce AuthMem-Bench, a controlled paired benchmark that holds the focal claim and downstream task fixed while varying only source authority. It evaluates write-time collapse, downstream authorization errors, and automatic authority preservation. Across seven consolidators based on widely used agent-memory systems and seven LLM backbones, we observe authority collapse in 48 of 49 evaluated configurations. In a controlled action-grounded evaluation, collapsed memories without authority metadata yield a mean unauthorized-action rate of 50.3%. In an end-to-end evaluation, automatically predicted and persisted authority labels reduce the observed unauthorized-action rate from 16.9% to 0.0%, while benign task success remains essentially unchanged. These findings show that memory-driven adaptation must preserve not only what was learned, but also the authority under which it may be reused.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.