알고부터 실천까지: 주식 시장에서 LLM 트레이딩 에이전트를 위한 메모리 기반 평가 기준
From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets
대규모 언어 모델(LLM) 에이전트가 자본 시장에서 수익을 창출할 수 있는지 평가하는 것은 점점 더 종단 간 트레이딩으로 정의되고 있습니다. 즉, 에이전트를 과거 시장에 배치하고 거래를 수행한 후 포트폴리오 수익률을 측정합니다. 이러한 방식은 두 가지 평가 오류에 취약합니다. 첫째, 장기간의 백테스팅은 종종 최첨단 LLM의 지식 제한과 겹쳐져, 에이전트가 기억하는 주식 티커, 날짜, 가격 및 시장 상황 정보가 투자 판단을 대체하게 됩니다. 둘째, 원시 수익률은 주식 선택 능력에 대한 정확한 지표가 아니기 때문에, 긍정적인 성과는 실제 알파(투자 수익)보다는 시장 베타, 투자 스타일 노출 또는 유리한 시장 환경으로 인해 발생할 수 있습니다. 저희는 이러한 문제를 해결하기 위해 KTD-Fin (Knowing-To-Doing Financial Benchmark, 알고부터 실천까지 금융 평가 기준)이라는 종단 간 주식 시장 트레이딩 평가 기준을 소개합니다. KTD-Fin은 데이터 측면의 마스킹 프로토콜을 사용하여 프롬프트와 도구를 통해 핵심 식별자 및 캘린더 정보를 일관되게 익명화하여, 과거 시장 기억과 투자 결정 과정을 분리합니다. 또한, 포트폴리오 수익률을 시장 베타, 스타일 노출 및 주식 선택 알파 구성 요소로 분해하는 Barra 스타일의 성과 분석 프레임워크를 포함하고 있습니다. 2024년부터 2026년까지 중국 CSI300 지수를 대상으로 평가된 10개의 최첨단 LLM 에이전트에 대한 실험 결과, 마스킹은 에이전트의 판단 과정을 크게 변화시키며, 익명화된 요인 기반 추론으로 이끌었습니다. 성과 분석 결과, 누출 제어 평가 하에서 LLM 에이전트의 총 수익률은 주로 수동적인 시장 및 스타일 노출에 의해 설명되며, 지속적인 주식 선택 알파를 보여주는 증거는 제한적이었습니다. 이러한 연구 결과는 금융 LLM 평가 기준이 단순히 에이전트가 돈을 벌는지 여부를 평가하는 것뿐만 아니라, 수익의 원천이 전송 가능한 투자 기술을 반영하는지 여부를 평가해야 함을 시사합니다. 저희는 KTD-Fin을 누출 제어 및 성과 분석 기반 LLM 트레이딩 에이전트 평가를 위한 재현 가능한 템플릿으로 공개합니다.
Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often overlap with the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to substitute for investment reasoning. Second, raw returns are a noisy proxy for stock-selection ability, since positive performance may come from market beta, style exposure, or favorable regimes rather than genuine alpha. We introduce KTD-Fin (Knowing-To-Doing Financial Benchmark), an end-to-end stock-market trading benchmark that addresses both issues. KTD-Fin uses a data-side masking protocol to anonymize key identifiers and calendar information consistently across prompts and tools, separating historical market memory from investment decision-making. It also incorporates a Barra-style performance attribution framework that decomposes portfolio returns into market, style, and stock-selection alpha components. Across ten frontier LLM agents evaluated on the Chinese CSI300 over a 2024--2026 window, masking substantially changes agent rationales, pushing them towards anonymized factor-based reasoning. Attribution analysis further shows that LLM agents' cumulative returns under leakage-controlled evaluation are largely explained by passive market and style exposure, with limited evidence of persistent stock-selection alpha. These findings suggest that financial LLM benchmarks should evaluate not only whether an agent makes money, but also whether the source of returns reflects transferable investment skill. We release KTD-Fin as a reproducible template for leakage-controlled and attribution-aware evaluation of LLM trading agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.