GAUGE: 정답이 없는 상황에서 에이전트가 구축한 금융 모델 평가
GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
금융 모델은 공개된 정보와 분석가의 가정을 결합하여 예측과 가치 평가를 수행합니다. 일부 구성 요소는 기계적으로 검증할 수 있지만, 예측값, 할인율, 목표 가격 등은 종종 여러 가지 합리적인 답을 제시할 수 있습니다. 기존의 벤치마크는 이러한 결과물을 단일 전문가 참조값을 기준으로 평가하는 경향이 있습니다. 동일한 회사에 대해 독립적으로 구축된 분석 모델을 사용하여 65개 회사의 108쌍을 비교한 결과, 중앙값 기준 점수는 0.33으로 나타났으며, 92.6%가 0.70 미만의 점수를 받았고, 동시 제작된 모델 쌍 중 어느 것도 암묵적 가격이 10% 이내로 일치하지 않았습니다. 따라서 허용 오차 기반의 평가는 전문가들 사이에서 이미 존재하는 의견 불일치를 벌점으로 처리할 수 있습니다. 우리는 GAUGE를 소개합니다. 이는 에이전트가 구축한 가치 평가 모델을 단일 정답이 아닌 실제 분석가의 관행에 비추어 평가하는 벤치마크입니다. GAUGE는 1,001개의 공급업체에서 분류한 분석 업무 자료와 196개 작업으로 구성된 평가 세트를 사용하며, 세 가지 수준의 실제 관행을 반영하고, 56가지 감사 항목, 8가지 유효성 검증 단계 및 결정론적 구조 검사를 포함합니다. 우리는 55명의 참가자가 참여한 그룹별 연구, 회사 기반 교차 검증 및 심사관 안정성 평가를 통해 이 벤치마크를 검증했습니다. '실패 인지 점수' $φ_0$에서 숙련된 분석가의 평균은 88.3점, 주니어는 66.0점, 금융학도들은 43.2점을 기록했습니다. 24개의 에이전트와 1,011개의 평가 결과물을 비교한 결과, 가장 높은 점수를 받은 에이전트는 53.4점으로 학생 평균보다 높았지만, 모든 숙련된 분석가 및 대부분의 주니어 분석가보다 낮았습니다. 이 에이전트는 기계적 요소에서 93%, 심사 요소에서 78%를 통과했으며, 전체 중앙값으로 26점의 격차가 있었습니다. 현재 에이전트는 모델 구축 능력은 뛰어나지만, 가치 평가 판단 능력은 상대적으로 부족합니다. 우리는 방법론, 익명화된 데이터 계층, 제한적인 학습 데이터 세트, 버전 관리된 48개 작업 평가 코어 및 보류 중인 업데이트 풀을 공개합니다.
Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $φ_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.