2606.16092v1 Jun 15, 2026 cs.CV

VinQA: 실세계 멀티모달 문서 질의응답을 위한 시각적 요소를 통합한 장문 답변 생성

VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA

Gyeonghun Kim
Gyeonghun Kim
Citations: 424
h-index: 5
Youngsoo Jang
Youngsoo Jang
Citations: 85
h-index: 4
Hyesoo Kong
Hyesoo Kong
Citations: 2
h-index: 1
Kyunghwan An
Kyunghwan An
Citations: 3
h-index: 1
Jae Sub Huh
Jae Sub Huh
Citations: 0
h-index: 0
Stanley Jungkyu Choi
Stanley Jungkyu Choi
Citations: 4
h-index: 2

실제 문서에서는 텍스트와 함께 표, 차트, 사진, 다이어그램 등 다양한 시각적 요소들이 복잡하게 배치되어 있습니다. 그러나 현재까지 멀티모달 대규모 언어 모델(MLLM)을 활용한 문서 질의응답 연구는 대부분 텍스트만으로 답변을 생성하여 이러한 시각적 요소들을 충분히 활용하지 못하고 있습니다. 본 논문에서는 VinQA라는 데이터셋을 소개합니다. 이 데이터셋은 장문 답변 생성을 위한 것이며, 인용된 시각적 요소들이 해당 요소를 뒷받침하는 텍스트와 함께 문서 페이지 내에 명시적으로 통합되어 있습니다. 이러한 작업을 지원하기 위해, 원본 문서 페이지 이미지를 MLLM에 입력하는 두 가지 인코딩 방법을 연구했습니다. 첫 번째 방법은 '페이지 인코딩'으로, 전체 페이지 이미지를 사용하여 시각적 요소의 경계 상자를 표시하고, 이 상자 영역을 인용 가능한 단위로 처리합니다. 두 번째 방법은 '모달리티 인코딩'으로, 각 페이지를 분석하여 텍스트와 시각적 요소를 분리하여 인코딩하고, 잘라낸 시각적 요소를 인용 가능한 단위로 사용합니다. 실험에서는 GroUSE를 확장한 M-GroSE라는 멀티모달 평가 프레임워크를 사용하여 답변의 완전성, 관련성, 충실성 및 불가능성을 4가지 차원에서 평가했습니다. 또한, 시각적 요소 인용 정확도를 직접 측정하기 위해 Visual Source F1을 사용했습니다. 비공개 최첨단 모델이 여전히 VinQA 테스트 데이터셋에서 가장 높은 점수를 기록하지만, 오픈 소스 Qwen2.5-VL 모델을 학습 데이터셋으로 미세 조정하면 성능이 크게 향상되어 이러한 격차를 줄일 수 있습니다. 모달리티 인코딩은 긴 텍스트, 많은 시각적 요소 및 다양한 인용 요구 사항을 가진 복잡한 문서에 대해 초기에는 더 강력한 성능을 보입니다. 그러나 VinQA 데이터셋으로 학습한 후에는 페이지 인코딩이 비슷한 수준의 성능을 달성하여, 모달리티 인코딩에서 사용되는 명시적인 분석 과정 없이도 효과적으로 경쟁할 수 있습니다. 마지막으로, MLLM 기반 평가 도구인 Visual G-Eval은 미세 조정된 모델들이 시각적 요소들을 의미적으로 적절한 위치에 삽입하고, 충실한 뒷받침 텍스트를 함께 제공한다는 것을 확인했습니다.

Original Abstract

Real-world documents combine text with tables, charts, photographs, and diagrams arranged in diverse layouts, yet existing research on multimodal large language models (MLLMs) for document QA predominantly produces text-only responses, underutilizing these visual elements. We introduce VinQA, a dataset for long-form answer generation where cited visual elements are explicitly interleaved with their supporting text and grounded in relevant document pages. To support this task, we study two encoding methods for feeding raw document page images into an MLLM, along with their visual-element citation mechanisms: (1) Page Encoding, which directly encodes full-page images with bounding boxes of visual elements and treats these boxed regions as citable units; and (2) Modality Encoding, which parses each page to extract text and crop visual elements, encodes them separately, and uses these cropped elements as citable units. In our experiments, we propose M-GroSE, a multimodal evaluation framework extending GroUSE to assess answers along four dimensions: completeness, answer relevancy, faithfulness, and unanswerability. We additionally report Visual Source F1 to directly measure visual citation accuracy. Although proprietary frontier models still achieve the best overall scores on the VinQA test split, fine-tuning open Qwen2.5-VL models on the training split substantially improves their performance and narrows this gap. Modality Encoding is initially more robust for complex documents with long text, many visual elements, and diverse citation requirements. After training on VinQA, however, Page Encoding reaches a comparable level, competing effectively even without the explicit parsing used in Modality Encoding. Finally, Visual G-Eval, an MLLM-based judge, confirms that fine-tuned models insert visual elements at semantically appropriate positions with faithful supporting text.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!