2606.10949v1 Jun 09, 2026 cs.AI

너무 잘 기억하는 것?: 메모리 강화 모델에서의 아첨 평가 및 완화

Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models

D. Bikel
D. Bikel
Citations: 20,646
h-index: 20
Aparna Balagopalan
Aparna Balagopalan
Citations: 804
h-index: 14
S. Bensal
S. Bensal
Citations: 26
h-index: 1
Axel Magnuson
Axel Magnuson
Citations: 19
h-index: 2

지속적인 메모리 시스템은 LLM이 사용자의 믿음을 시간 경과에 따라 저장함으로써 더욱 유용해질 수 있다고 약속합니다. 그러나 우리는 이러한 시스템이 모델의 정확성을 저하시키는 동시에, 모델이 정확성보다 사용자 동의를 우선시하는 아첨 현상을 체계적으로 증폭시킨다는 것을 보여줍니다. 본 연구에서는 이 효과를 체계적으로 평가하기 위해 MIST라는 벤치마크를 도입합니다. MIST는 과학, 의학 및 도덕적 추론 영역에서 사용자가 그럴듯한 오해를 표현하는 인공적으로 생성된 다중 대화 데이터셋입니다. 최첨단 메모리 시스템 세 가지와 모델 패밀리 다섯 개에 대한 실험 결과, 메모리가 모든 조건에서 아첨 행동을 증폭시키는 것으로 나타났으며, 일부 경우 in-context 기반 모델보다 아첨 비율이 최대 25배 더 높았습니다. 오류 분석 결과, 주요 원인은 메모리 추출 과정인 것으로 보입니다. 손실 압축 방식으로 분산된 조각으로 정보를 저장하는 과정에서 사용자의 오해는 그대로 유지되는 반면, 수정적인 맥락은 삭제됩니다. 이러한 결과를 바탕으로 우리는 아첨 현상을 크게 줄이면서 동시에 사실적 기억 능력은 기존 시스템과 동등하거나 뛰어넘는 두 가지 간단한 완화 방법을 제안합니다.

Original Abstract

Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by systematically amplifying sycophancy, wherein models prioritize agreement with users over accuracy. We conduct the first systematic evaluation of this effect, introducing MIST: a benchmark of synthetically generated multi-turn conversations where users express plausible misconceptions in scientific, medical, and moral reasoning domains. Testing across three state-of-the-art memory systems and five model families reveals that memory amplifies sycophantic behavior across all conditions, with up to 25x higher sycophancy rates than in-context baselines. Error analyses suggest memory extraction as the primary culprit: lossy compression into discrete snippets encodes user misconceptions while discarding corrective context. Based on these results, we propose two lightweight mitigations that substantially reduce sycophancy while matching or exceeding memory systems at factual recall.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!