2608.03292v1 Aug 04, 2026 cs.AI

DocTrace: 계층적 증거 그래프 추론을 통한 추적 가능한 긴 문서 시각 질의 응답

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

Hong Chen
Hong Chen
Citations: 23
h-index: 3
Lei Xiang
Lei Xiang
Citations: 42
h-index: 1
Zhicheng Guan
Zhicheng Guan
Citations: 0
h-index: 0
Xiaocong Lin
Xiaocong Lin
Citations: 0
h-index: 0
Zhenghua Lei
Zhenghua Lei
Citations: 0
h-index: 0
Teng Hu
Teng Hu
Citations: 129
h-index: 3
Bolei He
Bolei He
Citations: 11
h-index: 2
Long Zeng
Long Zeng
Citations: 0
h-index: 0

긴 문서 시각 질의 응답(LongDocVQA)은 멀티모달 거대 언어 모델(MLLM)이 여러 페이지에 분산된 다양한 문서 요소를 찾아 통합하고, 이를 바탕으로 추론하는 것을 요구합니다. 기존 접근 방식들은 엔드투엔드 MLLM, 검색 증강 생성(RAG) 파이프라인, 그리고 문서 에이전트 등을 포함하며, 이러한 방식들은 종종 추론 과정에서 어떻게 근거가 점진적으로 구성되는지를 명시적으로 표현하고 검증하는 메커니즘이 부족하여, 답변 정확도와 추적 가능성이 제한됩니다. 본 논문에서는 LongDocVQA를 암묵적인 답변 예측 문제로 보는 대신, 명시적인 증거 그래프 추론 문제로 재정의합니다. 이를 위해, 우리는 명시적인 근거 추적을 가능하게 하는 계층적 프레임워크인 DocTrace를 제안합니다. 이 프레임워크는 근거 위치 찾기, 구조화된 문서 파싱, 그리고 증거 그래프 추론을 점진적으로 수행합니다. 이러한 능력을 효과적으로 학습하기 위해, 우리는 두 단계의 훈련 프레임워크를 개발했습니다: 먼저, 지도 미세 조정(SFT)을 통해 근거 위치 찾기와 그래프 추론 능력을 초기화하고, 그 다음에는 특정 작업에 맞춘 그룹 상대 정책 최적화(GRPO)를 수행하며, 이를 위해 특별히 설계된 보상을 활용하여 이러한 능력을 더욱 향상시킵니다. MMLongBench-Doc, LongDocURL, 그리고 SlideVQA 데이터셋에 대한 광범위한 실험 결과는 DocTrace가 기존의 공개 소스 모델과 독점적인 MLLM을 모두 능가한다는 것을 보여줍니다. 특히, Qwen3-VL-8B-Instruct 모델을 기준으로 했을 때, DocTrace는 세 가지 벤치마크에서 각각 14.4%, 11.3%, 그리고 11.7%의 절대적인 성능 향상을 달성했습니다. 경쟁력 있는 성능 외에도, DocTrace는 명시적인 노드 수준의 출처를 가진 추적 가능한 증거 그래프를 구축하여, 긴 문서 이해에 대한 투명하고 검증 가능성을 제공합니다.

Original Abstract

Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages. Existing approaches, including end-to-end MLLMs, retrieval-augmented generation (RAG) pipelines, and document agents, often lack explicit mechanisms to represent and verify how grounded evidence is progressively composed during reasoning, limiting both answer accuracy and traceability. In this paper, we cast LongDocVQA as an explicit evidence graph reasoning problem rather than implicit answer prediction. To this end, we propose DocTrace, a hierarchical framework that progressively performs evidence localization, structured document parsing, and evidence graph reasoning to enable explicit evidence provenance. To effectively learn these capabilities, we develop a two-stage training framework: joint Supervised Fine-Tuning (SFT) first initializes evidence localization and graph reasoning abilities, followed by task-specific Group Relative Policy Optimization (GRPO) with dedicated rewards to further optimize these capabilities. Extensive experiments on MMLongBench-Doc, LongDocURL, and SlideVQA demonstrate that DocTrace consistently outperforms both existing open-source baselines and proprietary MLLMs. Compared with the Qwen3-VL-8B-Instruct backbone, DocTrace achieves absolute improvements of 14.4, 11.3, and 11.7 points on the three benchmarks, respectively. Beyond competitive performance, DocTrace constructs traceable evidence graphs with explicit node-level provenance, enabling transparent and verifiable reasoning for long document understanding.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!