2607.01071v1 Jul 01, 2026 cs.IR

MemSyco-Bench: 에이전트 메모리에서의 아첨(Sycophancy) 성능 측정

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Yunbo Tang
Yunbo Tang
Citations: 5
h-index: 2
Qinggang Zhang
Qinggang Zhang
Citations: 7
h-index: 2
Jinsong Su
Jinsong Su
Citations: 74
h-index: 3
Zhishang Xiang
Zhishang Xiang
Citations: 68
h-index: 3
Yujie Lin
Yujie Lin
Citations: 147
h-index: 7
Zerui Chen
Zerui Chen
Citations: 8
h-index: 2
Zhimin Wei
Zhimin Wei
Citations: 27
h-index: 3
Ruqin Ning
Ruqin Ning
Citations: 0
h-index: 0

메모리는 현대 LLM 기반 에이전트의 핵심 구성 요소로, 단일 회화 어시스턴트에서 장기 협력자로 발전하는 데 중요한 역할을 합니다. 그러나 메모리가 항상 유익한 것은 아닙니다. 검색된 기억은 종종 '아첨'이라는 심각한 문제를 야기하며, 이는 에이전트가 사실 정확성이나 객관적인 추론을 희생하고 사용자에게 지나치게 부합하도록 만듭니다. 이러한 새로운 위험에도 불구하고 기존의 메모리 벤치마크는 주로 기억이 올바르게 저장, 검색 또는 업데이트되는지 여부를 평가하는 데 초점을 맞추고 있으며, 검색된 기억이 하위 수준의 추론 및 의사 결정에 미치는 영향은 간과하고 있습니다. 이러한 격차를 해소하기 위해, 에이전트 시스템에서 메모리 기반 아첨을 평가하기 위한 포괄적인 벤치마크인 MemSyco-Bench를 제안합니다. MemSyco-Bench는 언제 기억이 의사 결정에 영향을 미쳐야 하는지, 그리고 어떤 유효한 기억이 사용되어야 하는지를 측정합니다. 구체적으로, 에이전트가 기억을 사실 증거로 거부할 수 있는지, 해당 적용 범위를 존중하는지, 기억과 객관적인 증거 간의 충돌을 해결하는지, 메모리 업데이트를 추적하는지, 그리고 유효한 기억을 개인화에 사용하는지를 평가하는 다섯 가지 작업을 포함합니다. 관련 모든 자료는 다음 GitHub 주소(https://github.com/XMUDeepLIT/MemSyco-Bench)에서 커뮤니티 구성원들을 위해 제공됩니다.

Original Abstract

Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning. Despite this emerging risk, existing memory benchmarks primarily evaluate whether memories are correctly stored, retrieved, or updated, while overlooking how retrieved memories influence downstream reasoning and decision-making. To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems. MemSyco-Bench measures when memory should influence a decision and how valid memory should be used. Specifically, it covers five tasks that assess whether agents can reject memory as factual evidence, respect its applicable scope, resolve conflicts between memory and objective evidence, track memory updates, and use valid memory for personalization. All related resources are collected for the community at https://github.com/XMUDeepLIT/MemSyco-Bench.

0 Citations
0 Influential
36.324746787308 Altmetric
0.0 Score
Original PDF
12

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!