메모리 전염(Memory Contagion): 에이전트 메모리를 통한 평가자 편향의 시간 경과에 따른 확산
Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory
대규모 언어 모델(LLM) 에이전트는 장기적인 일관성을 유지하기 위해 점점 더 기억 시스템에 의존하고 있습니다. 최근 연구에서는 에이전트의 기억이 지속적인 통합 과정에서 저하되는 것으로 나타났습니다. 그러나 기존 연구는 기억이 편향되지 않은 경험에서 파생된다고 가정합니다. 본 연구에서는 새로운 현상인 '메모리 전염(Memory Contagion)'을 밝혀내고 공식화했습니다. 메모리 전염은 에이전트의 기억을 통해 평가자 편향이 시간 경과에 따라 확산되는 현상을 의미합니다. 저희는 에이전트가 편향된 평가자에 의해 훈련되거나 안내받으면, 그 경험이 편향될 수 있다는 것을 보여줍니다. 이러한 경험들이 저장되고 통합되어 메모리에 저장되면, 동일한 메모리 저장소에서 정보를 검색하는 후속 에이전트로 편향이 전파됩니다. 특히, 통합 과정이 완벽(oracle)하더라도 이러한 현상이 발생합니다. 두 가지 유형의 편향(길이 선호, 권위 편향)과 네 가지 실험 단계에 걸쳐 다음과 같은 결과를 확인했습니다: (1) 메모리 전염은 완벽한 통합(oracle 조건)에서도 발생하며, 이는 편향된 입력이 전염의 충분조건임을 입증합니다. (2) 통합 과정은 편향 유형에 따라 상반되는 영향을 미칩니다. 즉, 길이 선호 편향을 효과적으로 감소시키는 반면, 권위 편향을 예비적으로 증폭시킵니다(단일 실행 추정). 이는 편향 유형에 따른 상호 작용 가능성을 시사합니다. (3) 안전한 임계값은 관찰되지 않았습니다. 오염율이 p=0.2로 매우 낮은 경우에도 편향 전파가 감지됩니다. 본 연구의 결과는 현재 에이전트 메모리 설계의 중요한 취약점을 드러내며, 시간 경과에 따른 편향 전파를 측정하기 위한 공식적인 도구를 제공합니다.
Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade during continuous consolidation. However, existing research assumes memories are derived from unbiased experiences. In this work, we identify and formalize a novel phenomenon: Memory Contagion -- the cross-temporal propagation of evaluator bias through agent memory. We show that when agents are trained or guided by biased evaluators, their experiences become biased; when these trajectories are stored and consolidated into memory, the bias propagates to future agents retrieving from the same memory store, even when consolidation is perfect (oracle). Across two bias types (length preference, authority bias) and four experimental phases, we demonstrate: (1) Memory Contagion occurs even with perfect consolidation (oracle condition), proving that biased input is a sufficient cause of contagion; (2) Consolidation has opposite effects depending on bias type -- robustly attenuating length bias while preliminarily amplifying authority bias (single-run estimate), suggesting a bias-type-dependent interaction; (3) No observed safe threshold: bias propagation is detected at contamination rates as low as p=0.2. Our findings expose a critical vulnerability in current agent memory designs and provide formal tools for measuring cross-temporal bias propagation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.