2607.29433v1 Jul 31, 2026 cs.CL

알고 실행하기: LLM 개인화에서 기억 활용에 대한 연구

Know It, Act on It: Investigating Memory Utilization in LLM Personalization

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Emmanuele Chersoni
Emmanuele Chersoni
Citations: 1,191
h-index: 20
Zhaoxin Feng
Zhaoxin Feng
Citations: 14
h-index: 2

대규모 언어 모델(LLM) 에이전트가 개인 맞춤형 동반자로 발전함에 따라, 메모리는 핵심 기능으로 부상했습니다. 그러나 LLM은 지식 활용 문제에 직면하는데, 맥락 내에 관련 사용자 선호도가 완전히 포함되어 있더라도 이를 바탕으로 행동하지 못하는 경우가 있습니다. 에이전트가 이전에 공유된 사용자 선호도가 중요해야 하는 상황에서 응답을 맞춤 설정하는 데 실패할 때, 모델이 해당 정보를 기억하지 못한 것인지, 아니면 기억했지만 사용하지 못한 것인지 불분명합니다. 이러한 문제를 분리하기 위해, 우리는 동일한 사용자 선호도에 대한 '알기(Know)' 및 '실행(Act)' 테스트를 결합한 독립적인 평가 패러다임을 도입했습니다. 16개의 시스템과 5가지 메모리 아키텍처에 대한 대규모 실험을 수행하여 표현 강도의 세 가지 수준으로 포함된 1,000개의 선호도를 평가했습니다. 그 결과, '알기'와 '실행'의 결과 사이에 큰 격차가 있음을 보여줍니다. 에이전트는 종종 사용자 선호도에 대한 회상 테스트를 통과하지만, 동일한 선호도가 관련 행동 시나리오에 반영되지 않는 경우가 많습니다. 메모리 아키텍처는 이러한 격차를 줄이지만, 특히 건강 및 치료와 관련된 선호도에서 활용도는 여전히 약하며, 이 경우 실행 실패가 가장 큰 현실적인 영향을 미칩니다.

Original Abstract

As large language model (LLM) agents evolve into personalized companions, memory has emerged as a core capability. However, LLMs face a knowledge utilization problem: they may fail to act on relevant user preferences even when they are fully present in context. When an agent fails to tailor its response in a context where previously shared user preferences should matter, it is unclear whether the model failed to remember that information or remembered it but failed to use it. To isolate this breakdown, we introduce a decoupled evaluation paradigm that administers paired Know and Act tests to the same user preference. We conduct large-scale experiments across 16 systems and five memory architectures, evaluating 1,000 preferences embedded at three levels of expression strength. Our results show a large gap between Know and Act outcomes: agents often pass the recall test for a user preference but fail to reflect that same preference in the paired behavioral scenario. While memory architectures reduce this gap, utilization remains especially weak for health and therapy-related preferences, where failures to act carry the greatest real-world stakes.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!