2608.04570v1 Aug 05, 2026 cs.CL

개인화의 환상: LLM이 사용자 프로필을 어떻게 허구로 만들어내는지, 그리고 자체 모니터링이 왜 오해를 불러일으키는가

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

Yushi Sun
Yushi Sun
Citations: 212
h-index: 6
Yanjie Zhang
Yanjie Zhang
Citations: 21
h-index: 2
Rui Sheng
Rui Sheng
Citations: 63
h-index: 6

지속적인 기억을 가진 개인화된 LLM이 점점 더 많이 사용되고 있지만, 이러한 모델들이 생성하는 사용자 모델의 정확성은 충분히 검토되지 않았습니다. 본 연구에서는 LLM이 증거로 뒷받침되지 않는 사용자 속성을 허구적으로 만들어내는 현상인 '과잉 추론(Over-Inference, OI)'을 분석합니다. 우리는 150개의 인물을 포함하는 MirageBench 데이터셋을 소개합니다. 이 데이터셋은 고정관념적인, 반고정관념적인, 그리고 중립적인 프로필로 균형을 이루며, '상상력 그레이디언트'를 포괄하는 6가지 개인화 작업을 포함합니다. 또한 독립적인 평가자가 판단한 정확성을 객관적인 기준으로 삼아 (400개의 주장에 대해 익명 인간 어노테이터와 비교하여 Cohen's kappa = 0.863, 이분법적 kappa = 0.900으로 검증), 7개 계열의 12개 모델에 대한 리더보드를 제공합니다 (총 143616개의 평가된 주장이 포함). 연구 결과, 과잉 추론은 광범위하게 나타났습니다. 12개 모델 모두에서 35%에서 49%의 주장이 과잉 추론되었으며 (모델 간 평균 41.6%, 주장 가중 평균 41.8%), 어떤 모델도 이 현상에서 자유롭지 않았습니다. 더욱 주목할 만한 점은 '자기 모니터링 역전(Self-Monitoring Inversion)' 현상을 발견했다는 것입니다. 즉, 모델 선택 단계에서 모델이 자체적으로 평가한 과잉 추론 정도는 독립적인 평가자가 측정한 과잉 추론 정도와 음의 상관 관계를 보였습니다 (rho = -0.60, p = 0.044; 탐색적, 넓은 부트스트랩 신뢰 구간 [-0.90, +0.06], n = 12). 과잉 추론이 가장 적다고 보고하는 모델일수록 실제로 가장 많은 허구를 만들어내는 경향을 보였으며, 따라서 모델 자체의 신뢰도 평가는 모델 비교를 위한 오해를 불러일으킬 수 있는 정보입니다. 하지만 단일 모델 내부에서는 자체 감사 결과가 해당 모델의 주장을 어느 정도 정확하게 평가하는 데 도움이 됩니다 (AUROC 0.58--0.83). 또한, 과잉 추론은 작업에 따라 달라지는 경향이 있으며 (27%--59%), 다중 라운드 테스트에서 추론된 속성은 거의 선형적으로 축적되며 수정 사항이 거의 없습니다. MirageBench는 모델 자체 보고가 아닌 외부 검증이 신뢰할 수 있는 개인화를 위한 더 안정적인 기반임을 보여줍니다.

Original Abstract

Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!