S

Sijia Chen

Total Citations
12
h-index
2
Papers
1

Publications

#1 2605.27887v1 May 27, 2026

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlations; they fail to evaluate the complete PM decision pipeline. We introduce PortBench, a benchmark spanning six heterogeneous asset classes over ten years. PortBench comprises two layers: a static QA dataset of 6,269 questions across seven task templates, and a dynamic five-stage allocation pipeline. To evaluate these layers, we introduce two metrics: a dual-layer correlation score for inter-class hedging and intra-class concentration, and CEPS, which quantifies how reasoning errors compound across pipeline stages. We further evaluate under three stress regimes and risk profiles, and support real-time evaluation to mitigate pretraining contamination on historical markets. Evaluating ten frontier LLMs, we find that despite strong financial QA performance, 90\% of model-profile cases fail to outperform equal-weight allocation in 2024, and this deficit persists across other market regimes; models that satisfy every procedural constraint still suffer large drawdowns under stress. Our source code is available at \href{https://github.com/AgenticFinLab/portbench}{this https URL}.

Ningxin Su Sijia Chen Yuxuan Zhao
2 Citations