2602.07642v1 Feb 07, 2026 cs.AI

멀티모달 대규모 언어 모델을 활용한 효율적인 표 검색 및 이해

Efficient Table Retrieval and Understanding with Multimodal Large Language Models

Haoyang Fang
Haoyang Fang
Citations: 201
h-index: 6
Boran Han
Boran Han
Citations: 356
h-index: 9
Bernie Wang
Bernie Wang
Citations: 80
h-index: 2
Cuixiong Hu
Cuixiong Hu
Citations: 176
h-index: 4
Shuai Zhang
Shuai Zhang
Amazon Web Services
Citations: 4,978
h-index: 26
Zhuoyan Xu
Zhuoyan Xu
Citations: 232
h-index: 6
Bonan Min
Bonan Min
Citations: 9
h-index: 2

표 데이터는 재무 보고서, 필기 기록, 문서 스캔 등 다양한 실제 시나리오에서 이미지 형태로 자주 등장합니다. 이러한 시각적 표현은 구조적 복잡성과 시각적 복잡성을 동시에 내포하고 있어 기계가 이해하는 데 독특한 어려움을 줍니다. 최근 멀티모달 대규모 언어 모델(MLLM)의 발전으로 표 이해 분야에서 유망한 결과가 나오고 있지만, 기존 연구들은 대개 관련된 표가 이미 주어져 있다고 가정합니다. 그러나 보다 실용적인 시나리오는 사용자 질의에 답하기 위해 대규모 컬렉션 내에서 관련 표를 식별하고 이를 바탕으로 추론하는 과정을 포함합니다. 이러한 간극을 해소하기 위해 본 논문에서는 MLLM이 대규모 표 이미지 컬렉션에 대한 질의에 답변할 수 있도록 하는 프레임워크인 TabRAG를 제안합니다. 우리의 접근 방식은 먼저 공동 학습된 시각-텍스트 파운데이션 모델을 사용하여 후보 표를 검색한 다음, MLLM을 활용하여 후보들을 정교하게 재순위화(reranking)하고, 마지막으로 선택된 표에 대해 MLLM이 추론하여 답변을 생성합니다. 48,504개의 고유한 표를 포함한 8개 벤치마크에 걸쳐 88,161개의 학습 샘플과 9,819개의 테스트 샘플로 구성된 새로운 데이터셋에 대한 광범위한 실험을 통해, 본 프레임워크가 기존 방법 대비 검색 재현율을 7.0%, 답변 정확도를 6.1% 크게 향상시킴을 입증하였으며, 이는 실제 표 이해 작업을 위한 실용적인 솔루션을 제공합니다.

Original Abstract

Tabular data is frequently captured in image form across a wide range of real-world scenarios such as financial reports, handwritten records, and document scans. These visual representations pose unique challenges for machine understanding, as they combine both structural and visual complexities. While recent advances in Multimodal Large Language Models (MLLMs) show promising results in table understanding, they typically assume the relevant table is readily available. However, a more practical scenario involves identifying and reasoning over relevant tables from large-scale collections to answer user queries. To address this gap, we propose TabRAG, a framework that enables MLLMs to answer queries over large collections of table images. Our approach first retrieves candidate tables using jointly trained visual-text foundation models, then leverages MLLMs to perform fine-grained reranking of these candidates, and finally employs MLLMs to reason over the selected tables for answer generation. Through extensive experiments on a newly constructed dataset comprising 88,161 training and 9,819 testing samples across 8 benchmarks with 48,504 unique tables, we demonstrate that our framework significantly outperforms existing methods by 7.0% in retrieval recall and 6.1% in answer accuracy, offering a practical solution for real-world table understanding tasks.

2 Citations
0 Influential
13 Altmetric
67.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!