2603.16557v1 Mar 17, 2026 cs.AI

BenchPreS: 지속적 메모리 LLM의 맥락 인지 개인화된 선호도 선택성을 위한 벤치마크

BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs

Albert No
Albert No
Citations: 91
h-index: 6
Sangyeon Yoon
Sangyeon Yoon
Citations: 43
h-index: 4
Hyesoo Hong
Hyesoo Hong
Citations: 13
h-index: 2
Wonje Jeung
Wonje Jeung
Citations: 143
h-index: 8
SunKyoung Kim
SunKyoung Kim
Citations: 445
h-index: 7
Yongil Kim
Yongil Kim
Citations: 156
h-index: 4
Heuiyeen Yeen
Heuiyeen Yeen
Citations: 127
h-index: 4
Wooseok Seo
Wooseok Seo
Citations: 12
h-index: 2

대규모 언어 모델(LLM)은 사용자 선호도를 지속적인 메모리에 저장하여 여러 상호 작용에서 개인화를 지원하는 경우가 점점 더 많아지고 있습니다. 그러나 사회적 및 제도적 규범에 따라 제약을 받는 타사 통신 환경에서는 일부 사용자 선호도가 적용하기에 적절하지 않을 수 있습니다. 본 논문에서는 BenchPreS를 소개하며, 이는 메모리에 저장된 사용자 선호도가 통신 맥락에 따라 적절하게 적용되는지 또는 억제되는지를 평가합니다. Misapplication Rate (MR) 및 Appropriate Application Rate (AAR)라는 두 가지 상호 보완적인 지표를 사용하여, 최첨단 LLM조차도 맥락에 민감하게 선호도를 적용하는 데 어려움을 겪는다는 것을 발견했습니다. 선호도 준수 능력이 더 강한 모델은 과도하게 적용되는 경향이 있으며, 추론 능력이나 프롬프트 기반 방어 기술도 이 문제를 완전히 해결하지 못합니다. 이러한 결과는 현재 LLM이 개인화된 선호도를 맥락에 따라 달라지는 규범적 신호가 아닌 전역적으로 적용 가능한 규칙으로 취급한다는 것을 시사합니다.

Original Abstract

Large language models (LLMs) increasingly store user preferences in persistent memory to support personalization across interactions. However, in third-party communication settings governed by social and institutional norms, some user preferences may be inappropriate to apply. We introduce BenchPreS, which evaluates whether memory-based user preferences are appropriately applied or suppressed across communication contexts. Using two complementary metrics, Misapplication Rate (MR) and Appropriate Application Rate (AAR), we find even frontier LLMs struggle to apply preferences in a context-sensitive manner. Models with stronger preference adherence exhibit higher rates of over-application, and neither reasoning capability nor prompt-based defenses fully resolve this issue. These results suggest current LLMs treat personalized preferences as globally enforceable rules rather than as context-dependent normative signals.

4 Citations
0 Influential
4 Altmetric
24.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!