2607.19867v1 Jul 22, 2026 cs.CL

FinMMEval 2026 Task 2 개요: 다국어 금융 단답형 질문 응답

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

Preslav Nakov
Preslav Nakov
Citations: 8,470
h-index: 49
R. Elbadry
R. Elbadry
Citations: 8
h-index: 2
D. Dimitrov
D. Dimitrov
Citations: 199
h-index: 8
Xueqing Peng
Xueqing Peng
Citations: 656
h-index: 13
Haolun Wu
Haolun Wu
Citations: 914
h-index: 11
Yankai Chen
Yankai Chen
Citations: 466
h-index: 10
Xue Liu
Xue Liu
Citations: 14
h-index: 2
Lingfei Qian
Lingfei Qian
Citations: 359
h-index: 10
Jimin Huang
Jimin Huang
Citations: 249
h-index: 8
Zhuohan Xie
Zhuohan Xie
Citations: 327
h-index: 9
Fan Zhang
Fan Zhang
Citations: 3
h-index: 1
Georgi N. Georgiev
Georgi N. Georgiev
Citations: 178
h-index: 5
Vanshikaa Jani
Vanshikaa Jani
Citations: 2
h-index: 1
Yuyang Dai
Yuyang Dai
Citations: 10
h-index: 2
Jiahui Geng
Jiahui Geng
Citations: 826
h-index: 13
Yuxia Wang
Yuxia Wang
Citations: 28
h-index: 3
Ivan Koychev
Ivan Koychev
Citations: 2,575
h-index: 26
Veselin Stoyanov
Veselin Stoyanov
Citations: 21
h-index: 3
Yu Chen
Yu Chen
Citations: 7
h-index: 1
Ye Yuan
Ye Yuan
Citations: 12
h-index: 2
M. Song
M. Song
Citations: 67
h-index: 2

FinMMEval 2026 Task 2는 다국어 자료를 활용한 금융 관련 단답형 질문 응답 시스템을 평가합니다. 각 최종 테스트 항목은 영어 질문과 함께 영어, 중국어, 일본어, 스페인어 및 그리스어로 작성된 재무 제표 및 뉴스 기사를 제공합니다. 참가 시스템은 각 항목에 대해 JSONL 형식으로 간결한 답변 하나를 제출해야 합니다. 최종 테스트 데이터는 256개의 항목으로 구성되며, 난이도는 쉬운 수준과 전문가 수준으로 균등하게 나뉘어 있습니다. 각 난이도별로 32개의 기업 보고서 그룹에 대한 네 가지 질문 템플릿이 사용됩니다. 정답은 제출 과정에서 공개되지 않으며, 시스템은 주최측이 보유한 참고 답변과의 ROUGE-1 F1 점수를 기준으로 순위가 매겨집니다. 최종 순위표에는 12개의 제출물이 포함되었습니다. 가장 우수한 시스템들은 성능이 매우 유사하며, 상위 네 개 시스템 간의 ROUGE-1 F1 점수 차이가 1% 미만입니다. 제출된 시스템 관련 논문에서는 검색 기반 생성(retrieval-augmented generation), 다국어 자료 처리, 구조화된 프롬프트 활용, 답변 압축 및 검증 전략 등이 설명되어 있습니다.

Original Abstract

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates instantiated over 32 company-report groups. Gold answers were withheld during submission, and systems were ranked by macro-averaged item-level ROUGE-1 F1 against organizer-held reference answers. The final leaderboard includes 12 ranked submissions. The strongest systems are closely clustered, with the top four separated by less than one percentage point in ROUGE-1 F1. The submitted system papers document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.

1 Citations
0 Influential
24.5 Altmetric
123.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!