2608.04095v1 Aug 04, 2026 cs.AI

FinPerMA: 이론 기반, 사건 중심의 LLM 에이전트를 위한 개인 맞춤형 기억 성능 평가 기준

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Chi Zhang
Chi Zhang
Citations: 24
h-index: 3
Kang Zhou
Kang Zhou
Citations: 75
h-index: 3

대규모 언어 모델(LLM) 에이전트는 금융 자문과 같은 고위험 영역에서 개인 비서로 점점 더 많이 사용되고 있지만, 이러한 에이전트가 장기간에 걸쳐 개별 사용자 모델을 유지하고 업데이트할 수 있는지 여부는 불분명합니다. 기존의 개인 맞춤형 기억 성능 평가 기준은 주로 사실 정보의 저장력을 테스트하거나, 제약 조건이 약한 모델 생성 경로에 의존하며, 사건 기반 선호도 적응에 대한 탐색이 부족했습니다. 본 연구에서는 사건을 중심으로 설계된 FinPerMA라는 새로운 평가 기준을 제시합니다. 이 기준은 고정된 장기 투자자 데이터를 사용하여 개인 맞춤형 기억 성능을 평가합니다. FinPerMA의 데이터 생성 파이프라인은 결정적인, 이론 기반 영향 규칙, 제어된 LLM 스토리텔링, 자동화된 품질 검사를 결합합니다. 또한, '충격(Shock)' 이후의 시점에서 에이전트가 중요한 사건을 지속적인 사용자 모델에 통합했는지 여부를 확인할 수 있는 체크포인트를 제공합니다. 276명의 가상 사용자에 대한 2,994개의 질문으로 구성된 데이터셋에서, 최첨단 LLM 7개와 최대 7가지의 메모리 설정을 사용하여 실험을 수행한 결과, 어떤 설정도 충분히 높은 성능을 보이지 않았습니다. 전체 정확도는 약 0.47, 객관식 문제에 대한 정확도는 약 39%였습니다. 분석 결과, 요약 기반 메모리는 사실 정보를 비교적 잘 유지하지만 개인 맞춤화에 필요한 선호도 신호를 손실하는 경향이 있습니다. 따라서 간단한 검색 방법이 특수 목적의 메모리 시스템보다 더 나은 성능을 보이는 경우가 있으며, 특히 충격 이후 이러한 격차는 더욱 벌어집니다.

Original Abstract

Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!