2605.27887v1 May 27, 2026 cs.AI

PortBench: 상관관계를 고려한 LLM 기반 포트폴리오 관리 전 과정 벤치마크

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

Ningxin Su
Ningxin Su
Citations: 41
h-index: 3
Sijia Chen
Sijia Chen
Citations: 12
h-index: 2
Yuxuan Zhao
Yuxuan Zhao
Citations: 1
h-index: 1

LLM(대규모 언어 모델)은 다양한 금융 분야에서 뛰어난 성능을 보여주었지만, 중요한 금융 의사 결정 과제인 포트폴리오 관리(PM)는 아직 제대로 된 벤치마크가 부족합니다. 기존 벤치마크는 두 가지 주요한 한계를 가지고 있습니다. 첫째, 자산 간의 상관 관계 구조를 고려하지 않아 진정으로 분산된 포트폴리오와 집중된 포트폴리오를 구별하기 어렵고, 둘째, 실제 시나리오에서 포괄적인 PM 의사 결정 과정을 평가하지 못합니다. 본 논문에서는 10년간의 데이터를 사용하여 6가지 이질적인 자산군을 다루는 벤치마크인 PortBench를 소개합니다. PortBench는 정적 질의응답 데이터셋과 동적 5단계 할당 파이프라인이라는 두 가지 상호 보완적인 계층으로 구성됩니다. 첫 번째 계층은 7가지 작업 템플릿에 걸쳐 6,269개의 상관 관계 기반 질문을 포함하는 정적 QA(질의응답) 데이터셋이며, 두 번째 계층은 전체 PM 의사 결정 주기를 반영하는 동적인 5단계 할당 파이프라인입니다. 이러한 계층을 평가하기 위해, 우리는 두 가지 전용 지표를 도입했습니다. 첫째, 제안된 포트폴리오가 자산 간의 헤징 효과를 활용하고 자산 내 집중화를 피하는지 측정하는 이중 계층 상관 관계 점수이고, 둘째, 추론 오류가 파이프라인 단계에 따라 어떻게 누적되는지를 정량화하는 CEPS(Compounded Errors in Portfolio Selection) 지표입니다. 또한, 우리는 세 가지 역사적인 스트레스 상황과 위험 프로필 하에서 전략의 견고성 및 투자자 일치성을 평가합니다. 10개의 최첨단 LLM을 평가한 결과, 정적 금융 QA에서는 뛰어난 성능을 보였지만, 모델-프로파일 조합의 90%가 기본적인 균등 가중 할당 방식보다 성능이 떨어졌으며, 모든 절차적 제약을 만족하는 모델조차도 스트레스 상황에서 심각한 손실(catastrophic drawdowns)을 경험했습니다. 본 연구의 소스 코드는 다음 링크에서 확인할 수 있습니다: [https://github.com/AgenticFinLab/portbench](this https URL).

Original Abstract

Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlations; they fail to evaluate the complete PM decision pipeline. We introduce PortBench, a benchmark spanning six heterogeneous asset classes over ten years. PortBench comprises two layers: a static QA dataset of 6,269 questions across seven task templates, and a dynamic five-stage allocation pipeline. To evaluate these layers, we introduce two metrics: a dual-layer correlation score for inter-class hedging and intra-class concentration, and CEPS, which quantifies how reasoning errors compound across pipeline stages. We further evaluate under three stress regimes and risk profiles, and support real-time evaluation to mitigate pretraining contamination on historical markets. Evaluating ten frontier LLMs, we find that despite strong financial QA performance, 90\% of model-profile cases fail to outperform equal-weight allocation in 2024, and this deficit persists across other market regimes; models that satisfy every procedural constraint still suffer large drawdowns under stress. Our source code is available at \href{https://github.com/AgenticFinLab/portbench}{this https URL}.

2 Citations
0 Influential
28.431471805599 Altmetric
11.0 Score
Original PDF
3

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!