FinanceHarness: 자율적인 금융 심층 연구 프레임워크
FinanceHarness: Autonomous Financial Deep Research Framework
LLM(대규모 언어 모델)과 자율 에이전트 기술의 발전으로, 심층 연구는 가장 널리 활용되는 에이전트 기반 제품 중 하나가 되었습니다. 그러나 대부분의 심층 연구 시스템은 범용 보고서를 작성하는데, 이는 금융 심층 연구에는 적합하지 않습니다. 금융 연구는 과거 패턴을 분석하고 미래 이벤트를 예측하기 위한 전문 지식을 요구합니다. 따라서 금융 심층 연구를 자동화하려면, 연구 에이전트를 제어하는 계층적인 구조와 미래 정보 유출을 방지하는 검증 가능하고 특정 시점의 벤치마크가 필요합니다. 본 논문에서는 FinanceHarness를 소개합니다. FinanceHarness는 금융 관련 도구 및 실무 전문가의 지침에 따른 워크플로우를 실행하여, 금융 심층 연구 프로세스 전체를 자동화합니다. 여기에는 환경 및 데이터 구축, 에이전트 실행 루프, 그리고 보상 모델링이 포함됩니다. 또한, 우리는 논문 주제 기반의 연구 질문과 평가 기준을 결합한 FinanceGym을 제안합니다. FinanceGym은 사전에 정의된 기준과 이후의 결과를 모두 고려하며, 전문가 검증 결과 82%의 합격률을 보입니다. 선도적인 LLM 및 에이전트조차도 평가 기준에서 40% 미만의 점수를 기록하여, FinanceGym이 높은 난이도를 가지고 있으며 상당한 개선 여지가 있음을 보여줍니다. 동일한 오픈 가중치 기반 모델을 사용하는 FinanceHarness는 전체 평가 기준 점수를 25.3%에서 32.4%로 향상시켰습니다. FinanceHarness는 https://github.com/Yijia-Xiao/FinanceHarness 에서 확인할 수 있습니다.
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%, demonstrating the effectiveness of our specialized harness design. However, even pairing FinanceHarness with the most cutting edge LLM (e.g. Opus-5), the FinanceGym score is below 45%, showing that it is a challenging benchmark for financial deep research. Leaderboard is available at: https://financegym.github.io/ and FinanceHarness code is available at: https://github.com/Yijia-Xiao/FinanceHarness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.