DRIFTLENS: 개인화된 언어 모델에서 기억에 의해 유발되는 추론 변화 측정
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
개인화는 모델이 사용자에게 제공하는 내용에 영향을 미치지만, 본 연구에서는 개인이 추론 과정에도 영향을 줄 수 있음을 보여줍니다. 최신 LLM은 사용자의 속성, 선호도 및 이전 컨텍스트를 저장하여 이를 향후 프롬프트에 반영함으로써 개인화된 상호 작용을 수행합니다. 본 연구는 이러한 기억이 정답이 없는 개방형 질문에서 추론 과정을 어떻게 변화시키는지 조사합니다. 이러한 효과를 정량화하기 위해, 우리는 참(ground-truth) 데이터 없이 각 추론 단계를 값 범주에 매핑하고, 특정 질문의 기억 기반 추론 경로와 그렇지 않은 추론 경로 간의 차이를 측정하는 DRIFTLENS라는 프레임워크를 제안합니다. 먼저, DRIFTLENS가 내용 없는 실용적인 노이즈와 의미 있는 추론 변화를 구별하는지 검증했습니다. 4개의 LLM과 연령, 직업, 장애 등 10가지 사용자 속성 범주에 대해, 사용자 속성 기반 기억은 각 모델의 실용적인 노이즈 수준 이상으로 중간에서 큰 추론 변화를 유발하며, 최종 답변이 유창하고 관련성이 높으며 타당성을 유지하는 경우에도 이러한 현상이 나타납니다. 다음으로, GRPO 및 DPO 기반의 사후 훈련 방법을 사용하여 추론 변화를 줄이는 효과를 평가했습니다. 두 방법 모두 추론 변화를 감소시키지만, 어느 한쪽이 항상 우월하지 않으며, 하위 작업 능력, 유용성 및 지시 따르기에 미치는 영향은 모델과 보상 함수에 따라 달라집니다. 이러한 결과는 기억 기반의 추론 변화가 개인화된 언어 모델에서 측정 가능한, 부분적으로 완화될 수 있는 오류 유형임을 시사합니다.
Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future prompts. We study whether such memory reshapes reasoning on open-ended questions where no single ground-truth answer exists. To quantify this effect, we introduce DRIFTLENS, a ground-truth-free framework that maps each expressed reasoning step to a value category and measures divergence between a question's no-memory trajectory and its trajectory under injected user-attribute memory. We first validate that DRIFTLENS distinguishes content-free pragmatic noise from substantive reasoning changes. Across four LLMs and 10 user-attribute categories, including age, occupation, and disability, user-attribute memory induces medium-to-large reasoning drift above each model's pragmatic-noise floor, even when final answers remain fluent, on-topic, and plausible. We then evaluate GRPO- and DPO-based post-training methods for reducing drift. Both reduce drift, but neither uniformly dominates; effects on downstream capability, helpfulness, and instruction following are model-and reward-dependent. These results suggest that memory-induced reasoning drift is a measurable and only partly mitigated failure mode of personalized language models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.