2608.04077v1 Aug 04, 2026 cs.AI

FinProBench: 전문가의 업무 결과물에서 파생된 역할 기반 평가 기준을 활용한 금융 AI 에이전트 평가

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Chi Zhang
Chi Zhang
Citations: 24
h-index: 3
Kang Zhou
Kang Zhou
Citations: 75
h-index: 3

금융 AI 에이전트를 평가하기 위해서는 실제 전문적인 업무와 관련된 기준이 필요합니다. 기존의 평가 방법은 주로 작업 지시문이나 모델 출력 결과를 바탕으로 기준을 설정하지만, 실무자의 결과물에서만 드러나는 암묵적인 기준은 간과하는 경향이 있습니다. 본 논문에서는 전문적인 금융 업무를 위한 벤치마크인 FinProBench와 재사용 가능한 파이프라인인 역할 기반 평가 기준 생성(Role-Grounded Rubric Construction, RGRC)을 소개합니다. RGRC는 실무자가 생산한 결과물을 활용하여 역할을 기준으로 평가 기준을 도출하는 과정을 포함하며, 총 4단계로 구성됩니다: 결과물 수집, 역량 추출, 평가 기준 합성 및 검증. 이 방법은 암묵적인 기준을 반영하고, 품질 수준을 구별하며, 동일 역할 내의 다양한 작업에 적용될 수 있습니다. 분석 전에 우리는 57가지 직업을 결과물 유형에 따라 30개의 풍부한 사전 지식을 가진 일반적인 역할과 27개의 제한된 사전 지식을 가진 전문화된 역할로 분류했습니다. 모든 역할에서 프롬프트만 사용하는 방법은 일반적인 역할에서는 RGRC와 거의 유사한 성능(89.2% vs. 90.7%)을 보이지만, 전문화된 역할에서는 RGRC가 훨씬 더 우수한 성능(99.1% vs. 78.0%)을 나타냅니다. 이러한 결과는 모델의 사전 지식에 일반적인 규칙이 잘 표현되어 있을 때 프롬프트 엔지니어링만으로도 유사한 평가 기준을 만들 수 있지만, 사전 지식 범위를 넘어서는 전문적인 기준은 실무 기반 접근 방식이 필수적임을 시사합니다. FinProBench는 57가지 직업, 8개의 금융 하위 산업 및 161가지 결과물 유형에 걸쳐 총 1,723건의 선별된 결과물을 기반으로 구축되었으며, 초기 평가 세트로는 7개 하위 산업 내 20가지 역할에 대한 20개의 완전한 작업으로 구성됩니다. 다양한 LLM 평가 모델과 역할 수준의 평가 기준을 활용하여, 인간이 수행한 결과물이 평균적으로 가장 높은 순위를 차지합니다(100점 만점에 73.7점, 다른 시스템은 각각 70.3점, 70.2점, 69.6점). 모든 시스템은 95% 신뢰 구간 내에서 중복되는 경향을 보이며 상호 보완적인 강점을 가집니다. 역할 수준에서 평가 기준을 재사용하면 각 평가 기준을 처음부터 작성하는 데 필요한 예상 작업량을 6.7배 줄일 수 있습니다.

Original Abstract

Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables. We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics from deliverables produced by practitioners in the same role. RGRC comprises four stages: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation. Its rubrics capture tacit standards, distinguish quality levels, and transfer across tasks within a role. Before analysis, we classified 57 occupations by deliverable genre into 30 prior-rich conventional roles and 27 prior-sparse role-specialized roles. Across all roles, Prompt-only nearly matches RGRC for conventional roles (89.2% vs. 90.7%), but RGRC substantially outperforms it for role-specialized roles (99.1% vs. 78.0%). This split indicates that prompt engineering can approximate rubrics when conventions are well represented in model priors, while professional grounding is essential for standards beyond those priors. FinProBench is built from 1,723 curated deliverables spanning 57 occupations, 8 financial sub-industries, and 161 deliverable types, and releases an initial evaluation set of 20 complete tasks covering 20 roles in 7 sub-industries. With heterogeneous LLM judges and role-level rubrics, human deliverables rank first on average (73.7 vs. 70.3, 70.2, and 69.6 out of 100), while all four systems show overlapping 95% confidence intervals and complementary strengths. Reusing rubrics at the role level reduces estimated per-task construction effort by 6.7 times relative to authoring each rubric from scratch.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!