2602.15958v1 Feb 17, 2026 cs.CL

DocSplit: 문서 패킷 인식 및 분할을 위한 종합적인 벤치마크 데이터셋 및 평가 방법

DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting

Md Mofijul Islam
Md Mofijul Islam
Citations: 7
h-index: 1
Md Sirajus Salekin
Md Sirajus Salekin
Amazon Web Services
Citations: 523
h-index: 14
N. Balakrishnan
N. Balakrishnan
Citations: 3
h-index: 1
Niharika Jain
Niharika Jain
Citations: 5
h-index: 1
Spencer Romo
Spencer Romo
Citations: 3
h-index: 1
Bob Strahan
Bob Strahan
Citations: 2
h-index: 1
Diego A. Socolinsky
Diego A. Socolinsky
Citations: 1,708
h-index: 16
Boyi Xie
Boyi Xie
Citations: 2,596
h-index: 7
Vincil Bishop
Vincil Bishop
Citations: 4
h-index: 1

실제 응용 분야에서 문서 이해는 종종 여러 문서가 연결된 이종의 다중 페이지 문서 패킷을 처리해야 하는 경우가 많습니다. 시각적 문서 이해 분야에서 최근의 발전에도 불구하고, 문서 패킷을 개별 단위로 분리하는 기본적인 문서 패킷 분할 작업은 아직까지 충분히 다루어지지 않았습니다. 본 연구에서는 최초의 종합적인 벤치마크 데이터셋인 DocSplit을 제시하며, 대규모 언어 모델의 문서 패킷 분할 능력을 평가하기 위한 새로운 평가 지표를 함께 소개합니다. DocSplit은 다양한 복잡성을 가진 다섯 개의 데이터셋으로 구성되어 있으며, 다양한 문서 유형, 레이아웃 및 다중 모드 설정을 포함합니다. 본 연구에서는 문서 경계를 식별하고, 문서 유형을 분류하며, 문서 패킷 내에서 정확한 페이지 순서를 유지해야 하는 DocSplit 작업을 정의합니다. 벤치마크는 페이지 순서가 뒤바뀐 경우, 문서가 섞인 경우, 그리고 명확한 구분 없이 구성된 문서와 같은 실제적인 문제들을 다룹니다. 저희는 다양한 다중 모드 LLM에 대한 광범위한 실험을 수행하여, 현재 모델이 복잡한 문서 분할 작업을 처리하는 능력에서 상당한 성능 격차가 있음을 확인했습니다. DocSplit 벤치마크 데이터셋과 제안된 새로운 평가 지표는 법률, 금융, 의료 및 기타 문서 중심 분야에서 필수적인 문서 이해 능력을 향상시키기 위한 체계적인 프레임워크를 제공합니다. 본 연구에서는 향후 문서 패킷 처리 연구를 촉진하기 위해 데이터셋을 공개합니다.

Original Abstract

Document understanding in real-world applications often requires processing heterogeneous, multi-page document packets containing multiple documents stitched together. Despite recent advances in visual document understanding, the fundamental task of document packet splitting, which involves separating a document packet into individual units, remains largely unaddressed. We present the first comprehensive benchmark dataset, DocSplit, along with novel evaluation metrics for assessing the document packet splitting capabilities of large language models. DocSplit comprises five datasets of varying complexity, covering diverse document types, layouts, and multimodal settings. We formalize the DocSplit task, which requires models to identify document boundaries, classify document types, and maintain correct page ordering within a document packet. The benchmark addresses real-world challenges, including out-of-order pages, interleaved documents, and documents lacking clear demarcations. We conduct extensive experiments evaluating multimodal LLMs on our datasets, revealing significant performance gaps in current models' ability to handle complex document splitting tasks. The DocSplit benchmark datasets and proposed novel evaluation metrics provide a systematic framework for advancing document understanding capabilities essential for legal, financial, healthcare, and other document-intensive domains. We release the datasets to facilitate future research in document packet processing.

1 Citations
0 Influential
8 Altmetric
41.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!