2608.06146v1 Aug 06, 2026 cs.AI

PaDoc: 문서 분석을 위한 레이아웃 기반 병렬 디코딩

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Chun Yuan
Chun Yuan
Citations: 51
h-index: 4
Chen Li
Chen Li
Citations: 67
h-index: 6
Jing Lyu
Jing Lyu
Citations: 14
h-index: 2
Hao Yu
Hao Yu
Citations: 24
h-index: 3
Jiabo Zhan
Jiabo Zhan
Citations: 8
h-index: 1
Dongxu Yue
Dongxu Yue
Citations: 0
h-index: 0
Chong Sun
Chong Sun
Citations: 74
h-index: 5
Linnan Zhao
Linnan Zhao
Citations: 0
h-index: 0
Rui Chen
Rui Chen
Citations: 0
h-index: 0
Jinglin Wang
Jinglin Wang
Citations: 16
h-index: 2
Kang Liu
Kang Liu
Citations: 0
h-index: 0

종단 간 문서 파서는 통합 인터페이스를 제공하지만, 페이지 레이아웃과 영역 콘텐츠를 하나의 자기 회귀 시퀀스로 직렬화합니다. 이러한 방식은 독립적인 영역들을 전체 내용의 길이에 따라 증가하는 디코딩 경로에 의존하게 만들며, 반면 크롭 기반 2단계 파서들은 영역 수준의 병렬성을 제공하지만, 반복적인 시각적 사전 채우기와 단편화된 페이지 컨텍스트라는 단점을 가집니다. 본 논문에서는 전체 페이지 컨텍스트를 유지하면서 의존성을 제거하기 위해, 예측된 레이아웃을 공유된 페이지 표현에 대한 분기 구조로 취급하는 레이아웃 기반 파서인 PaDoc을 제안합니다. 영역의 충분성 가정 하에서, 우리는 레이아웃 스트림과 영역 콘텐츠 분기가 동시에 진행되는 접두사 조건부 인수분해를 도출하며, 이를 통해 디코딩 깊이를 가장 긴 레이아웃-콘텐츠 경로로 줄입니다. 이러한 인수분해는 단일 MLLM 내에서 구현되며, 패킹된 가변 길이 상위 어텐션은 표준 다음 토큰 훈련 하에서 가시성을 유지하고, 마스크된 병렬 디코딩을 통해 평가되는 vLLM 백엔드가 동시 요청으로 처리하며 캐시에 저장된 공통 접두사를 재사용합니다. OmniDocBench Full 데이터셋에서 PaDoc은 전체 레이아웃 F1 점수 91.1을 달성했으며, 종단 간 파서 중 최고 수준인 전체 점수 94.24를 기록했습니다. 또한 Text Edit (0.038) 및 Formula CDM (95.59)에서 최상의 성능을 보였습니다. 384페이지의 하위 집합과 하나의 A800 GPU를 사용하여, PaDoc은 5개의 병렬 처리 수준에서 가장 빠른 종단 간 파서이며, 유효한 페이지 처리량을 67.4%에서 118%까지 향상시키고 P95 지연 시간을 39.2%에서 54.9%까지 줄였습니다. 코드 및 관련 정보는 다음 링크에서 확인할 수 있습니다: https://github.com/Longin-Yu/Padoc

Original Abstract

End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we propose PaDoc, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region-sufficiency assumption, we derive a prefix-conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout-content path. We realize this factorization within a single MLLM: packed variable-length ancestor attention preserves the visibility under standard next-token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache-resident shared-prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end-to-end parsers, a top-tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384-page subset and one A800 GPU, it is the fastest end-to-end parser at five concurrency levels, improving valid-page throughput by 67.4-118% and reducing P95 latency by 39.2-54.9% relative to a same-backbone Sequential SFT baseline. Code is available at https://github.com/Longin-Yu/Padoc

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!