2608.06108v1 Aug 06, 2026 cs.AI

대규모 언어 모델의 투자 논리 평가: 개인 맞춤형 금융 에이전트를 위한 실제 환경 기반 성능 측정

Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

Shuai Jia
Shuai Jia
Citations: 26
h-index: 2
Yuanhong Jiang
Yuanhong Jiang
Citations: 0
h-index: 0
Jingjie Zou
Jingjie Zou
Citations: 0
h-index: 0
Zhenghong Lin
Zhenghong Lin
Citations: 0
h-index: 0
Xusheng Yu
Xusheng Yu
Citations: 0
h-index: 0
Shijie Dai
Shijie Dai
Citations: 0
h-index: 0
Qiqi Huang
Qiqi Huang
Citations: 9
h-index: 2

투자 역량은 본질적으로 개인화되어 있습니다. 동일한 시장 정보라도 투자 목표, 기간, 포트폴리오 및 위험 감수 수준이 다른 투자자에게는 서로 다른 행동을 정당화할 수 있습니다. 그러나 현재 금융 분야의 대규모 언어 모델(LLM)은 정적인 질문-답변 방식이나 최종 수익과 손실만을 기준으로 평가됩니다. 전자는 에이전시(agency, 주체성)를 간과하고, 후자는 수익성이 좋은 행동이 얼마나 근거에 기반했는지, 투자자 프로필과 일치하는지, 아니면 단순히 운이 좋았는지 여부를 파악할 수 없습니다. 우리는 커뮤니티가 중요한 에이전트의 성능을 평가하기 위해 적절하지 않은 척도를 사용하고 있는지 질문합니다. 저희는 151명의 실제 투자자로부터 수집된 201,247건의 의사 결정 데이터를 포함하는 프로세스 중심의 벤치마크인 extsc{InvestLogicBench}를 소개합니다. 각 시나리오는 투자자의 extit{프로필}, 관찰 가능한 시장 extit{이벤트}, 투자 extit{추론}, 실행 가능한 extit{결정}, 그리고 지연된 extit{결과}의 순서로 구성됩니다. 벤치마크 데이터는 프로필 생성, 특정 시점 이벤트 매핑, 체계적인 논리, 투자 기간, 결과 및 사후 분석 정보를 포함하며, 이해, 프로필 기반 생성 및 전체 재현 기능을 지원합니다. 네 가지 주요 LLM을 대상으로 평가한 결과, 논리적 타당성은 평균 4/5점으로 비교적 높았지만, 이벤트의 근거는 0.8~2.8/5점으로 낮게 나타났습니다. 또한 수익률과 프로세스 품질도 서로 일치하지 않았습니다. 이러한 결과는 최종 결과만을 평가할 때 숨겨지는, 정교하지만 근거가 부족한 추론을 보여줍니다. 우리는 더 나아가 P$ ightarrow$E$ ightarrow$R$ ightarrow$D$ ightarrow$O(프로필-이벤트-추론-결정-결과) 모델이 데이터-시스템 인터페이스로서 기능해야 하며, 이를 위해 버전 관리된 프로필, 시간적 출처 정보, 검토 가능한 검색 결과, 의사 결정 기록 및 재현 가능한 결과를 제공해야 한다고 주장합니다. 금융 분야는 개인 맞춤형이고 중요한 영향을 미치는 에이전트에 대한 보다 광범위한 테스트 환경으로 활용될 수 있습니다.

Original Abstract

Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \textsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} trace: investor \textit{Profile}, observable market \textit{Events}, investment \textit{Reasoning}, executable \textit{Decision}, and delayed \textit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!