2607.24273v1 Jul 27, 2026 cs.CL

INS-ActBench: 대규모 언어 모델의 전문적인 계리 능력 평가를 위한 종합 벤치마크

INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

Changyu Chen
Changyu Chen
Citations: 106
h-index: 2
Chenwei Lin
Chenwei Lin
Citations: 14
h-index: 2
X. Xu
X. Xu
Citations: 4
h-index: 1

대규모 언어 모델(LLM)은 금융 추론 분야에서 강력한 잠재력을 보여주고 있지만, 기존 벤치마크는 종종 도메인 지식, 수리적 추론, 긴 문맥 이해 및 도구 사용 능력을 개별적으로 평가합니다. 이는 감사 가능하고, 문맥에 기반하며, 도구를 활용할 수 있는 실제적인 전문 업무 흐름을 평가하는 데 한계를 갖습니다. 본 연구에서는 대규모 언어 모델의 전문적인 계리 능력을 평가하기 위한 종합 벤치마크인 **INS-ActBench**를 소개합니다. INS-ActBench는 16개의 계리 협회에서 공개한 공무 시험 문제와 샘플 질문에서 추출한 12,050개의 질의응답 쌍으로 구성되어 있습니다. 이는 세 가지 하위 집합으로 나뉩니다: 표준화된 계리 지식을 평가하는 **INS-Act-Know**, 긴 문맥 기반 보험 사례 추론을 평가하는 **INS-Act-Case**, 그리고 검증 가능한 수치 결과를 생성하는 스프레드시트 및 R 코드 작업을 평가하는 **INS-Act-Practice**입니다. 9개의 대표적인 대규모 언어 모델과 인간 계리 전문가를 대상으로 한 실험 결과, 명확한 능력 차이가 나타났습니다. 최첨단 LLM은 표준화된 지식에 강점을 보이지만, 사례 추론, 도구 기반 워크플로우 및 관할 구역에 민감한 실무에서는 여전히 취약합니다. INS-ActBench는 신뢰할 수 있는 전문 지원을 위한 계리 LLM 개발의 재현 가능한 기반을 제공합니다. 코드 및 관련 정보는 https://github.com/FDU-INS/INS-ActBench 에서 확인할 수 있습니다.

Original Abstract

Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limits their ability to assess realistic professional workflows that require auditable, context-grounded, and tool-executable decisions. We introduce \textbf{INS-ActBench}, a comprehensive benchmark for evaluating professional actuarial capability in LLMs. INS-ActBench contains 12,050 Q\&A pairs from public exams and sample questions released by 16 actuarial associations. It covers three subsets: \textbf{INS-Act-Know} for standardized actuarial knowledge, \textbf{INS-Act-Case} for long-context insurance case reasoning, and \textbf{INS-Act-Practice} for spreadsheet and R-code tasks with verifiable numerical outputs. Experiments on nine representative LLMs and human actuarial experts reveal a clear capability boundary: frontier LLMs perform strongly on standardized knowledge, but remain much weaker in case reasoning, tool-based workflows, and jurisdiction-sensitive practice. INS-ActBench provides a reproducible foundation for developing actuarial LLMs toward reliable professional assistance. The code is available at https://github.com/FDU-INS/INS-ActBench.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!