2606.31522v1 Jun 30, 2026 cs.CL

FinPersona-Bench: 자율 금융 에이전트의 장기적인 심리 측정 안정성을 위한 벤치마크

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Preslav Nakov
Preslav Nakov
Citations: 8,470
h-index: 49
R. Elbadry
R. Elbadry
Citations: 8
h-index: 2
Xueqing Peng
Xueqing Peng
Citations: 656
h-index: 13
Yankai Chen
Yankai Chen
Citations: 466
h-index: 10
Xue Liu
Xue Liu
Citations: 14
h-index: 2
Zhuohan Xie
Zhuohan Xie
Citations: 327
h-index: 9
Fan Zhang
Fan Zhang
Citations: 16
h-index: 3
Ayesha Gull
Ayesha Gull
Citations: 3
h-index: 1
Muhammad Usman Safder
Muhammad Usman Safder
Citations: 9
h-index: 2

대규모 언어 모델(LLM)은 '자본 보존' 또는 '투기적 투자 회피'와 같은 명시적인 행동 지침을 통해 모든 의사 결정을 통제하도록 설계된 자율 금융 에이전트로 점점 더 많이 사용되고 있습니다. 그러나 실제로는 시장 상황이 장기간에 걸쳐 축적됨에 따라 이러한 지침들이 점차적으로 행동에 미치는 영향력이 감소하는 현상이 발생하는데, 이를 우리는 '명령 중요도 감소(MSD)'라고 정의합니다. MSD를 객관적으로 측정하기 위해, 우리는 FinPersona-Bench라는 시뮬레이션 벤치마크를 소개합니다. 이 벤치마크에서는 합성 시장에서 관찰 가능한 가격과 숨겨진 근본 가치를 분리하여, 세 가지 실패 모드에 대한 검증 가능한 평가가 가능합니다: 안정적인 시장에서의 신호 없이 거래, 급락 시의 공포 매도, 투기적 거품 발생 시의 근본 가치 무시. 우리는 18개의 선도적인 최첨단 및 오픈 소스 LLM을 평가했습니다. 각 모델은 엄격한 자본 보존부터 공격적인 성장까지 세 가지 행동 프로필 중 하나를 할당받았습니다. 분석 결과, MSD는 시간이 지남에 따라 누적되며 모델에 따라 다릅니다. 급락 시나리오에서는 정적 에이전트와 주기적인 명확화 과정을 거치는 에이전트 간의 행동 격차가 시뮬레이션의 첫 번째 분기에서 마지막 분기로 이동하면서 4.4배 증가합니다. 명확화 과정의 효과는 균일하게 긍정적이지 않습니다. 보수적인 에이전트는 신호가 낮은 시장에서 일관되게 도움을 받는 반면, 동일한 환경에서 공격적인 에이전트에게는 오히려 행동을 악화시킵니다. 이러한 결과는 안정적인 장기 운영을 위해서는 에이전트 프로필과 시장 상황에 따른 선택적이고 명확화 지침 기반 재정렬이 필요하다는 것을 시사합니다.

Original Abstract

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment. In practice, however, as market context accumulates over long horizons, these mandates gradually lose their behavioral influence, a phenomenon we formalize as Mandate Salience Decay (MSD). To measure MSD objectively, we introduce FinPersona-Bench, a simulation benchmark in which a synthetic market decouples observable price from hidden fundamental value, enabling falsifiable evaluation across three failure modes: trading without signal in calm markets, panic-selling during crashes, and ignoring fundamental value during speculative bubbles. Evaluating 18 leading frontier and open-source LLMs, each assigned one of three behavioral profiles ranging from strict capital preservation to aggressive growth, shows that MSD compounds over time and is model-dependent. In crash scenarios, the behavioral gap between static agents and those receiving periodic mandate re-grounding grows 4.4x from the first to the final quarter of the simulation. The effects of mandate re-grounding are not uniformly positive: it consistently helps conservative agents in low-signal markets but actively worsens behavior for aggressive agents in the same setting. These findings suggest that reliable long-horizon deployment requires selective, mandate-aware re-grounding based on agent profile and market regime.

1 Citations
0 Influential
24.5 Altmetric
123.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!