2606.05749v1 Jun 04, 2026 cs.CL

MARDoc: 다중 모드 긴 문서 질의응답을 위한 메모리 기반 정제 에이전트 프레임워크

MARDoc: A Memory-Aware Refinement Agent Framework for Multimodal Long Document QA

Hongtao Liu
Hongtao Liu
Citations: 95
h-index: 5
Jian Yang
Jian Yang
Citations: 14
h-index: 1
Qiyao Peng
Qiyao Peng
Citations: 108
h-index: 4
Kaifeng Chen
Kaifeng Chen
Citations: 201
h-index: 4
Yongqiang Liu
Yongqiang Liu
Citations: 28
h-index: 3
Xiaochen Zhang
Xiaochen Zhang
Citations: 0
h-index: 0
Qing Yang
Qing Yang
Citations: 74
h-index: 3

최근, 반복적인 검색-추론 에이전트는 다중 모드 긴 문서 질의응답 분야에서 상당한 잠재력을 보여주었습니다. 그러나 대부분의 기존 시스템은 검색 기록, 관찰 내용 및 중간 추론을 혼합하는 단일, 확장되는 컨텍스트를 유지합니다. 상호 작용이 누적됨에 따라 중요한 증거가 흩어지고 희석되어 다단계 추론 과정에서 노이즈가 발생합니다. 본 논문에서는 MARDoc이라는 메모리 기반 정제 에이전트 프레임워크를 제안합니다. MARDoc은 긴 문서 질의응답을 세 가지 특화된 에이전트로 분리합니다: 다중 수준의 다중 모드 검색을 수행하는 탐색기(Explorer), 상호 작용 기록을 구조화된 증거 및 추론 메모리로 정제하는 정제기(Refiner), 그리고 증거의 충분성을 검사하고 목표 지향적인 피드백을 제공하는 반사기(Reflector). 각 에이전트는 전체 누적된 상호 작용 기록 대신 동적으로 업데이트되는 구조화된 메모리에 의존합니다. 이러한 설계는 컨텍스트 노이즈를 줄이면서 답변에 중요한 사실과 그들의 논리적 종속성을 보존합니다. MMLongBench-Doc 및 DocBench 데이터셋에서의 실험 결과, MARDoc은 동일한 기반 모델을 사용하는 기존 시스템보다 뛰어난 성능을 보여주었으며, 구조화된 메모리가 에이전트 기반 문서 질의응답에 효과적임을 입증했습니다.

Original Abstract

Iterative retrieval-reasoning agents have recently shown promise for multimodal long-document question answering. However, most existing systems maintain a single growing context that mixes retrieval traces, observations, and intermediate reasoning. As interactions accumulate, key evidence becomes scattered and diluted, making multi-hop reasoning noisy. We propose MARDoc, a Memory-Aware Refinement Agent framework that decouples long-document QA into three specialized agents: an Explorer for multi-granularity multimodal retrieval, a Refiner for distilling interaction traces into structured evidence and reasoning memories, and a Reflector for checking evidence sufficiency and providing targeted feedback. Across iterations, the agents rely on a dynamically updated structured memory rather than a full accumulated interaction history. This design reduces context noise while preserving answer-critical facts and their logical dependencies. Experiments on MMLongBench-Doc and DocBench show that MARDoc achieves strong results, outperforming same-backbone baselines and demonstrating the effectiveness of structured memory for agentic document QA.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!