2608.04374v1 Aug 05, 2026 cs.CL

FinReportBench: 기관 수준의 금융 보고서 생성 성능 측정 및 개선

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

Wei Chen
Wei Chen
Citations: 400
h-index: 5
Jun Zhou
Jun Zhou
Citations: 24
h-index: 3
Ying Tang
Ying Tang
Citations: 56
h-index: 4
Xiaolu Zhang
Xiaolu Zhang
Citations: 52
h-index: 4
Yiyao Wang
Yiyao Wang
Citations: 37
h-index: 3
Zhenwei Tan
Zhenwei Tan
Citations: 0
h-index: 0
Wanli Gu
Wanli Gu
Citations: 59
h-index: 3

대규모 언어 모델은 유창한 금융 분석을 생성할 수 있지만, 유창성만으로는 보고서가 기관에 적합한지 판단할 수 없습니다. 본 연구에서는 전문가의 의견을 반영하여 기관 수준의 금융 보고서 생성 성능을 측정하고 개선하기 위한 벤치마크인 FinReportBench를 소개합니다. 전문가 검토 결과, 보고서의 내용 충실도, 기관 관련 요소 포함 여부, 출처 분야 및 시각적 표현 방식에서 반복적으로 나타나는 문제점이 발견되었습니다. 우리는 전문가의 판단 순서를 기반으로 다중 모드 증거 및 의사 결정 경계 검토를 통해 35개의 평가 항목을 구성하여 보고서의 활용 가능성, 내용 충실도 및 기관 관련 요소 완전성을 평가합니다. 10,000개의 균형 잡힌 중국어 및 영어 금융 연구 자료를 기반으로 세 가지 연구 주제와 두 가지 입력 수준에서 총 244개의 이중 언어 작업 세트를 구성했습니다. 각 작업은 공개된 질문, 재구성된 연구 경로 및 숨겨진 소스 데이터를 포함합니다. 세 개의 독립적인 평가 그룹이 전문가의 판단 순서를 거의 완벽하게 재현하여, 명확하고 관찰 가능한 기준이 신뢰할 수 있는 평가를 지원한다는 것을 보여줍니다. 9개의 모델 패밀리를 비교한 결과, 기본적인 활용 가능성은 거의 최고 수준에 도달했지만, 보고서 내용 충실도와 기관 관련 요소 완전성은 여전히 주요 개선 영역입니다. 모델 간 가장 큰 차이는 기본적인 보고서 구조보다는 생성 과정 제어, 정보 밀도 및 데이터 관리 측면에서 나타났습니다. 그런 다음, 벤치마크 기반의 기술 증류를 통해 반복적으로 발생하는 문제점을 재사용 가능한 생성 및 자체 검토 규칙으로 변환했습니다. 5개의 모델 패밀리에 대해 개발된 기술은 평균 G1 점수를 33.85점, 평균 G2 점수를 13.83점 향상시켰으며, 모든 쌍에서 G0 점수는 유지되었습니다. 코드 및 벤치마크 관련 자료는 https://github.com/MisterBrookT/finreportbench 에서 확인할 수 있습니다.

Original Abstract

Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery. We introduce FinReportBench, an expert-grounded benchmark for measuring and improving institution-grade financial report generation. Expert review reveals recurring gaps in report identity, institutional components, source discipline, and visual delivery. We derive a 35-item rubric through expert partial orders, multimodal evidence, and audits of decision boundaries, covering deliverability, report identity, and institutional completeness. Starting from 10,000 balanced Chinese and English financial-research source records, we curate 244 bilingual tasks across three research objects and two input tiers. Each task separates the public query, reconstructed research trajectory, and hidden source packet. Three independent judge families reproduce the expert partial order at near-ceiling rates, showing that bounded, observable criteria support reliable evaluation. Across nine model families, basic deliverability is nearly saturated, while report identity and institutional completeness remain the primary bottlenecks. The largest cross-model gaps concern generation-trace control, information density, and data discipline rather than basic report framing. We then use benchmark-guided skill distillation to turn recurrent failures into reusable generation and self-review constraints. Across five model families, the evolved skill improves mean G1 by 33.85 points and mean G2 by 13.83 points over paired no-skill runs while preserving G0 for every pair. Code and benchmark artifacts are available at https://github.com/MisterBrookT/finreportbench.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!