2608.04424v1 Aug 05, 2026 cs.CV

기준점(Anchor)을 활용한 사고: 문서 이해를 위한 실용적이고 효율적인 추론

Thinking with Anchors: Grounded and Efficient Document Reasoning

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Yuchen Zhu
Yuchen Zhu
Citations: 222
h-index: 11
Molei Tao
Molei Tao
Citations: 247
h-index: 11
Yujun Cai
Yujun Cai
Citations: 3,844
h-index: 21
Yongxin Chen
Yongxin Chen
Citations: 6
h-index: 2
Wenzhuo Xu
Wenzhuo Xu
Citations: 0
h-index: 0
Jason Kuen
Jason Kuen
Citations: 1,565
h-index: 18
Wanrong Zhu
Wanrong Zhu
Citations: 166
h-index: 7
Bing Shuai
Bing Shuai
Citations: 9,269
h-index: 25
Qin Zhang
Qin Zhang
Citations: 54
h-index: 2
Shilong Liu
Shilong Liu
Citations: 157
h-index: 2
Jiuxiang Gu
Jiuxiang Gu
Citations: 162
h-index: 6
Jing Shi
Jing Shi
Citations: 0
h-index: 0
Sichen Zhu
Sichen Zhu
Citations: 209
h-index: 5
Xuan Shen
Xuan Shen
Citations: 214
h-index: 8
Quanyi Wang
Quanyi Wang
Citations: 31
h-index: 3

기존의 문서 이해 성능 평가에서는 주로 페이지 요소의 위치 파악에 초점을 맞추었지만, 실제 환경에서의 문서 지능은 모델이 영역 의미, 공간 관계 및 시각적 구조에 대해 종합적으로 추론할 수 있어야 합니다. 본 논문에서는 ADOPD 2026을 제시합니다. 이는 ADOPD를 확장하여 페이지 분해를 통해 공간적으로 연계된 문서 이해를 가능하게 하는 추론 중심의 데이터셋입니다. ADOPD 2026은 ADOPD 2024 데이터셋에서 상속받은 페이지 기준점(anchor)에 사람이 직접 작성한 설명, 의미 태그 및 문서 영역에 기반한 체인 오브 소트(Chain-of-Thought, CoT) 추론 과정을 추가했습니다. 기존 방식이 사각형, 마스크 및 태그를 독립적인 지도 신호로 취급하는 반면, 우리는 텍스트 블록, 시각적 요소, 의미 레이블, 경계 상자 및 다각형 마스크를 문서 이해를 위한 공유된 시각적 기준점 어휘로 정의했습니다. 이러한 표현 방식은 세 가지 주요 기능을 지원합니다. 첫째, 영역 수준의 의미 태깅을 통해 모델은 페이지 컨텍스트와 지역적인 특징을 모두 활용하여 문서 요소 유형을 식별하도록 하며, 이를 통해 기존 레이아웃 성능 평가에서 숨겨지는 장기적인 의미 오류를 드러냅니다. 둘째, 통합된 시각-언어 연동 방식은 텍스트 영역과 시각적 요소를 좌표 또는 다각형 윤곽선과 함께 생성하여 검출 및 분할 결과를 구조화된 기준점으로 변환함으로써 다운스트림 추론 시스템에서 재사용할 수 있도록 합니다. 셋째, ADOPD 2026에서 파생된 DocCount 성능 평가 지표에 대한 현재 최첨단 모델의 성능은 여전히 제한적이며, 이는 문서 의미 이해를 위한 '기준점을 활용한 사고' 파이프라인의 필요성을 강조합니다. ADOPD 2026은 페이지 분해와 검증 가능한 시각적 기준점 추론을 연결하여 문서 이해를 단순 위치 파악에서 벗어나 기준점에 기반한 문서 지능으로 발전시키는 프레임워크를 제공합니다.

Original Abstract

Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region semantics, spatial relations, and visual structure. We present ADOPD 2026, a reasoning-oriented extension of ADOPD that turns page decomposition into spatially grounded document understanding. ADOPD 2026 enriches page anchors inherited from ADOPD 2024 dataset with human-cleaned captions, semantic tags, and generated chain-of-thought (CoT) traces grounded to document regions. Instead of treating boxes, masks, and tags as independent supervision signals, we cast text blocks, visual entities, semantic labels, bounding boxes, and polygon masks as a shared vocabulary of visual anchors. This representation supports three connected capabilities. First, region-level semantic tagging asks models to identify document element types from both page context and local appearance, revealing long-tail semantic failures that standard layout benchmarks often hide. Second, unified vision-language grounding generates text regions and visual entities together with coordinates or polygonal outlines, transforming detection and segmentation outputs into structured anchors that can be reused by downstream reasoning systems. Third, current state-of-the-art models still struggle with dense counting tasks evaluated on DocCount, a benchmark derived from ADOPD 2026, highlighting the need for the Thinking-with-Anchors pipeline in document semantic understanding. By connecting page decomposition to verifiable visual-anchor reasoning, ADOPD 2026 provides a task framework that moves document understanding beyond localization toward anchor-grounded document intelligence.

0 Citations
0 Influential
12.5 Altmetric
62.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!