2607.19856v1 Jul 22, 2026 cs.CL

FinMMEval 2026 Task 1 개요: 다국어 금융 객관식 질문 응답

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

Preslav Nakov
Preslav Nakov
Citations: 8,470
h-index: 49
R. Elbadry
R. Elbadry
Citations: 8
h-index: 2
D. Dimitrov
D. Dimitrov
Citations: 199
h-index: 8
Xueqing Peng
Xueqing Peng
Citations: 656
h-index: 13
Haolun Wu
Haolun Wu
Citations: 914
h-index: 11
Yankai Chen
Yankai Chen
Citations: 466
h-index: 10
Xue Liu
Xue Liu
Citations: 14
h-index: 2
Lingfei Qian
Lingfei Qian
Citations: 359
h-index: 10
Jimin Huang
Jimin Huang
Citations: 249
h-index: 8
Zhuohan Xie
Zhuohan Xie
Citations: 327
h-index: 9
Fan Zhang
Fan Zhang
Citations: 3
h-index: 1
Georgi N. Georgiev
Georgi N. Georgiev
Citations: 178
h-index: 5
Vanshikaa Jani
Vanshikaa Jani
Citations: 2
h-index: 1
Yuyang Dai
Yuyang Dai
Citations: 10
h-index: 2
Jiahui Geng
Jiahui Geng
Citations: 826
h-index: 13
Yuxia Wang
Yuxia Wang
Citations: 28
h-index: 3
Ivan Koychev
Ivan Koychev
Citations: 2,575
h-index: 26
Veselin Stoyanov
Veselin Stoyanov
Citations: 21
h-index: 3
Yu Chen
Yu Chen
Citations: 7
h-index: 1
Ye Yuan
Ye Yuan
Citations: 12
h-index: 2
M. Song
M. Song
Citations: 67
h-index: 2

FinMMEval 2026 Task 1은 영어, 중국어, 아랍어 및 힌디어로 구성된 다국어 금융 객관식 질문 응답 시스템의 성능을 평가합니다. 이 과제는 시스템이 전문 용어, 수치 해석 및 개념적 금융 추론과 관련된 금융 질문에 대해 올바른 답변을 선택할 수 있는지 테스트합니다. 최종 테스트 데이터 세트는 800개의 질문으로 구성되며, 각 언어당 200개의 질문이 포함됩니다. 정답은 제출 시 공개되지 않았으며, 각 언어는 정확도를 기준으로 독립적으로 순위가 매겨졌습니다. 최종 리더보드에는 영어 13개, 중국어 11개, 아랍어 11개, 힌디어 10개의 시스템이 포함되어 있으며, 이들 시스템은 정확도 순으로 나열되었습니다. 최고 정확도는 힌디어에서 92.0%, 영어 및 아랍어에서 97.5%이며, 최상위 팀들은 모든 언어에서 상위권에 속했습니다. 문서화된 시스템에서는 검색 증강, 직접 답변-옵션 점수 매기기, 언어별 프롬프트 사용, 선택적 자기 일관성, 신뢰도 검사 및 LLM 기반 검토 단계를 활용했습니다.

Original Abstract

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per language; gold answers were withheld during submission, and each language was ranked independently by accuracy. The final leaderboards contain 13 English, 11 Chinese, 11 Arabic, and 10 Hindi ranked submissions. Top accuracies range from 92.0% in Hindi to 97.5% in English and Arabic, with the same leading teams appearing near the top across all four languages. The documented systems used retrieval augmentation, direct answer-option scoring, language-specific prompting, selective self-consistency, confidence checks, and LLM-based review stages.

1 Citations
0 Influential
24.5 Altmetric
123.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!