TABVERSE: LLM 및 VLM에서 교차 형식 테이블 이해 성능 평가
TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs
대규모 언어 모델(LLM)과 시각-언어 모델(VLM)은 점점 더 많은 테이블 추론 작업에 사용되고 있지만, 테이블 표현 방식의 역할은 아직 충분히 연구되지 않았습니다. 실제로 동일한 테이블 내용은 HTML, Markdown, LaTeX와 같은 다양한 구조적 형식으로 나타나거나 렌더링된 이미지로 제공될 수 있습니다. 그러나 기존 평가에서는 종종 내용, 형식, 레이아웃 및 모달리티가 동시에 변하여 표현 방식의 영향을 분리하기 어렵습니다. 본 연구에서는 TABVERSE라는 제어된 다중 모드 테이블 벤치마크를 소개합니다. 이 벤치마크는 동일한 테이블 내용을 여러 구조적 형식과 렌더링된 이미지로 일치시키고, 질문 유형 및 난이도 태그를 포함하여 설계되었습니다. 이러한 설계는 테이블 내용이 고정된 상태에서 표현 방식의 영향을 체계적으로 평가할 수 있도록 합니다. 우리는 LLM과 VLM을 질의 응답(QA), 구조 이해 능력(SUC) 및 구조 재구성(SR)이라는 세 가지 작업에 대해 평가했습니다. 결과는 표현 방식 선택이 테이블 이해 성능에 상당한 영향을 미친다는 것을 보여줍니다. 모델은 일반적으로 렌더링된 이미지보다 구조화된 텍스트 형식으로 더 나은 성능을 보이지만, 이러한 차이는 작업, 모델 및 형식에 따라 달라집니다. HTML은 종종 가장 안정적인 텍스트 형식이며, 행에 민감한 구조적 작업과 구문적으로 사용 가능한 LaTeX 재구성은 여전히 어려운 과제로 남아 있습니다. 이러한 결과는 테이블 표현 방식이 신뢰할 수 있는 테이블 평가에서 중요한 요소임을 보여줍니다.
Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly evaluated on table reasoning tasks, but the role of table representation remains under-explored. In practice, the same table content may appear in different structural formats, such as HTML, Markdown, and LaTeX, or as rendered images. However, existing evaluations often let content, format, layout, and modality vary together, making it difficult to isolate representation effects. We introduce TABVERSE, a controlled multimodal table benchmark that aligns the same table content across multiple structural formats and rendered images, with question category and difficulty tags. This design enables systematic evaluation of representation effects while holding table content fixed. We evaluate LLMs and VLMs across three tasks: Question Answering (QA), Structural Understanding Capability (SUC), and Structure Reconstruction (SR). Our results show that representation choice substantially affects table understanding. Models generally perform better with structured text than with rendered images, but the size of this gap depends on the task, model, and format. HTML is often the most robust text format, while row-sensitive structural tasks and syntactically usable LaTeX reconstruction remain challenging. These findings show that table representation is a key factor in reliable table evaluation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.