2605.03903v2 May 05, 2026 cs.CL

CC-OCR V2: 실제 환경의 시각 문서 이해에서 LMM 실패 원인에 대한 세밀한 분석

CC-OCR V2: Fine-Grained Attribution of LMM Failures in Real-World Visual Document Understanding

Shuai Bai
Shuai Bai
Citations: 21,355
h-index: 20
Dayiheng Liu
Dayiheng Liu
Alibaba Group
Citations: 5,173
h-index: 18
Zhenghao Liu
Zhenghao Liu
Citations: 226
h-index: 9
Zulong Chen
Zulong Chen
Citations: 23
h-index: 2
Chunyi Peng
Chunyi Peng
Citations: 24
h-index: 3
Zhipeng Xu
Zhipeng Xu
Citations: 78
h-index: 5
Zhibo Yang
Zhibo Yang
Citations: 8,292
h-index: 26
Junhao Ji
Junhao Ji
Citations: 0
h-index: 0
Qing Liu
Qing Liu
Citations: 11
h-index: 1
Z. Qin
Z. Qin
Citations: 14
h-index: 2
Zebo Xu
Zebo Xu
Citations: 48
h-index: 4
Jianqiang Wan
Jianqiang Wan
Citations: 6,594
h-index: 10
Jun Tang
Jun Tang
Citations: 5,288
h-index: 7

최근 대규모 다중 모드 모델(LMM)은 OCR 중심의 문서 이해 및 처리 작업에서 놀라운 발전을 이루었습니다. 기존 벤치마크는 다양한 작업 전반에 걸쳐 LMM을 평가하여 실제 문서 처리 워크플로우를 반영하거나 문서 특성이 모델 성능에 미치는 영향을 분석합니다. 그러나 이러한 벤치마크는 조명, 화면 표시, 이미지 품질 및 촬영 방법과 같이 실제 환경의 문서 수집 조건에서 LMM의 신뢰성에 대한 제한적인 통찰력을 제공합니다. 이러한 격차를 해소하기 위해, 우리는 실제 문서 처리에서 LMM 실패 원인을 분석하는 포괄적인 벤치마크인 CC-OCR v2를 제시합니다. CC-OCR v2는 16개의 하위 작업, 74개의 응용 시나리오 및 7,093개의 샘플을 통해 문서 처리의 다섯 가지 핵심 기능을 평가하는 통합 프레임워크를 제공합니다. 이 벤치마크는 인식 영역에서 다섯 가지 평가 트랙, 열 가지 문서 카테고리 및 32개 언어를 포함합니다. 작업 수준의 평가 외에도 각 샘플은 세밀하게 정의된 열 가지 문서 요인에 대해 주석이 달려 있어 문서 유형 및 수집 조건 전반에 걸쳐 모델 실패를 체계적으로 분석할 수 있습니다. 17개의 대표적인 LMM을 대상으로 수행한 광범위한 실험 결과, 작업, 문서 카테고리 및 실제 환경 조건에 따라 상당한 성능 변동이 나타났습니다. 또한 전체 정확도가 비슷한 모델이라도 특정 문서 요인 하에서 근본적으로 다른 실패 패턴을 보이는 경우가 많으며, 이는 벤치마크 수준의 성능과 실제 응용 분야에서의 신뢰성 있는 배포 간의 중요한 격차를 보여줍니다. 데이터셋 및 평가 도구는 https://github.com/eioss/CC-OCR-V2 에서 공개적으로 이용할 수 있습니다.

Original Abstract

Recent Large Multimodal Models (LMMs) have achieved remarkable progress on OCR-centric document understanding and processing tasks. Existing benchmarks primarily evaluate LMMs across diverse tasks to reflect practical document-processing workflows or analyze how document characteristics influence model performance. However, they provide limited insight into the reliability of LMMs under real-world document acquisition conditions, where factors such as lighting, screen displays, imaging quality, and capture methods can substantially affect performance. To bridge this gap, we present CC-OCR v2, a comprehensive benchmark for attributing LMM failures in real-world document processing. CC-OCR v2 provides a unified evaluation framework covering five core document-processing capabilities through 16 subtasks, 74 application scenarios, and 7,093 samples. The benchmark spans five evaluation tracks, ten document categories, and 32 languages in the recognition suite. Beyond task-level evaluation, each sample is annotated with ten fine-grained document factors, enabling systematic attribution of model failures across document types and acquisition conditions. Extensive experiments on 17 representative LMMs reveal substantial performance variation across tasks, document categories, and real-world conditions. Moreover, models with comparable overall accuracy often exhibit fundamentally different failure patterns under specific document factors, highlighting a significant gap between benchmark-level performance and reliable deployment in practical applications. The dataset and evaluation toolkit are publicly available at https://github.com/eioss/CC-OCR-V2.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!