DocSplit: 문서 패킷 인식 및 분할을 위한 종합적인 벤치마크 데이터셋 및 평가 방법
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
실제 응용 분야에서 문서 이해는 종종 여러 문서가 연결된 이종의 다중 페이지 문서 패킷을 처리해야 하는 경우가 많습니다. 시각적 문서 이해 분야에서 최근의 발전에도 불구하고, 문서 패킷을 개별 단위로 분리하는 기본적인 문서 패킷 분할 작업은 아직까지 충분히 다루어지지 않았습니다. 본 연구에서는 최초의 종합적인 벤치마크 데이터셋인 DocSplit을 제시하며, 대규모 언어 모델의 문서 패킷 분할 능력을 평가하기 위한 새로운 평가 지표를 함께 소개합니다. DocSplit은 다양한 복잡성을 가진 다섯 개의 데이터셋으로 구성되어 있으며, 다양한 문서 유형, 레이아웃 및 다중 모드 설정을 포함합니다. 본 연구에서는 문서 경계를 식별하고, 문서 유형을 분류하며, 문서 패킷 내에서 정확한 페이지 순서를 유지해야 하는 DocSplit 작업을 정의합니다. 벤치마크는 페이지 순서가 뒤바뀐 경우, 문서가 섞인 경우, 그리고 명확한 구분 없이 구성된 문서와 같은 실제적인 문제들을 다룹니다. 저희는 다양한 다중 모드 LLM에 대한 광범위한 실험을 수행하여, 현재 모델이 복잡한 문서 분할 작업을 처리하는 능력에서 상당한 성능 격차가 있음을 확인했습니다. DocSplit 벤치마크 데이터셋과 제안된 새로운 평가 지표는 법률, 금융, 의료 및 기타 문서 중심 분야에서 필수적인 문서 이해 능력을 향상시키기 위한 체계적인 프레임워크를 제공합니다. 본 연구에서는 향후 문서 패킷 처리 연구를 촉진하기 위해 데이터셋을 공개합니다.
Document understanding in real-world applications often requires processing heterogeneous, multi-page document packets containing multiple documents stitched together. Despite recent advances in visual document understanding, the fundamental task of document packet splitting, which involves separating a document packet into individual units, remains largely unaddressed. We present the first comprehensive benchmark dataset, DocSplit, along with novel evaluation metrics for assessing the document packet splitting capabilities of large language models. DocSplit comprises five datasets of varying complexity, covering diverse document types, layouts, and multimodal settings. We formalize the DocSplit task, which requires models to identify document boundaries, classify document types, and maintain correct page ordering within a document packet. The benchmark addresses real-world challenges, including out-of-order pages, interleaved documents, and documents lacking clear demarcations. We conduct extensive experiments evaluating multimodal LLMs on our datasets, revealing significant performance gaps in current models' ability to handle complex document splitting tasks. The DocSplit benchmark datasets and proposed novel evaluation metrics provide a systematic framework for advancing document understanding capabilities essential for legal, financial, healthcare, and other document-intensive domains. We release the datasets to facilitate future research in document packet processing.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.