2605.29462v1 May 28, 2026 cs.CV

CFMME를 활용한 대규모 시각-언어 모델 성능 평가: 중국 금융 분야의 종합적인 다중 모드 평가 데이터셋

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

Qianben Chen
Qianben Chen
Citations: 152
h-index: 5
Xianyin Zhang
Xianyin Zhang
Citations: 97
h-index: 3
Lifan Guo
Lifan Guo
Citations: 69
h-index: 4
Feng Chen
Feng Chen
Citations: 39
h-index: 3
Chi Zhang
Chi Zhang
Citations: 24
h-index: 3
Yanzhi Liu
Yanzhi Liu
Citations: 57
h-index: 3

대규모 시각-언어 모델(LVLM)의 등장은 텍스트 기반 이해를 넘어 모델의 능력을 크게 확장하여, 시각 및 텍스트 양쪽의 정보를 통합적으로 활용하고 다양한 실제 응용 분야를 지원하게 되었습니다. 본 연구에서는 중국 내 금융 업무 전반에 걸쳐 LVLM의 인식, 이해, 추론 및 인지 능력에 대한 종합적인 평가를 수행하기 위해, 새로운 중국 금융 다중 모드 평가 데이터셋인 CFMME를 제안합니다. CFMME는 기초 학문 지식부터 복잡한 실제 응용까지 6,052개의 샘플로 구성되어 있으며, 8가지 주요 금융 이미지 모드와 4가지 핵심 다중 모드 작업을 포함합니다. 본 연구에서는 CFMME를 사용하여 대표적인 LVLM들을 철저하게 평가했습니다. 그 결과, 최첨단 모델은 질의 응답 작업에서 전반적으로 66.11%의 정확도를 달성했으며, 탐지, 인식 및 정보 추출 작업에서는 평균 점수가 77.18점으로 나타났습니다. 이는 현재 LVLM의 성능 개선 가능성이 매우 높다는 것을 시사합니다. 또한, 오류 원인, 모드 간 상호 작용 능력 및 다양한 환경 설정에 대한 상세 분석을 수행하여 향후 연구에 유용한 정보를 제공합니다. 본 연구에서 제시하는 CFMME가 LVLM, 특히 금융 분야의 다중 모드 작업 성능 개선에 기여할 수 있기를 기대합니다.

Original Abstract

The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual and textual modalities and supporting a broader range of real-world applications. To comprehensively evaluate the perception, understanding, reasoning, and cognition capabilities of LVLMs throughout the entire financial business workflow in Chinese contexts, we introduce CFMME, a novel Chinese financial multimodal evaluation benchmark. CFMME comprises 6,052 instances spanning from fundamental academic knowledge to complex real-world applications, covering eight primary financial image modalities and four core multimodal tasks. On CFMME, we conduct a thorough evaluation of representative LVLMs. The results show that the state-of-the-art model attains an overall accuracy of 66.11\% on the question answering task and an average score of 77.18 on the detection, recognition, and information extraction tasks, indicating substantial room for improvement in current LVLMs. In addition, we conduct detailed analyses of error causes, cross-modal capabilities, and multi-orientation settings, yielding valuable insights for future research. We hope that CFMME will spur further progress in LVLMs, especially by improving their performance on multiple multimodal tasks in the financial domain.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!