다중 모드 LLM을 신뢰할 수 있는 차트 데이터 추출기로 만드는 방법: 벤치마크 및 학습 프레임워크
Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework
차트 데이터 추출은 차트 이미지로부터 데이터 테이블을 역추적하는 기술로, 재현성, 분석, 검색 및 재설계를 위해 필수적입니다. 기존의 인터랙티브 도구는 신뢰성이 높지만 번거롭고, 혼합형 시스템은 효율적이지만 일반화 가능성이 부족합니다. 최근 개발된 다중 모드 대규모 언어 모델(MLLM)은 차트 해석을 위한 통합 인터페이스를 제공하지만, 특히 데이터 레이블이 없는 경우 정확한 데이터 테이블 추출 능력은 불분명합니다. 본 연구에서는 이러한 능력을 평가하기 위해 다양한 실제 차트를 포함하는 벤치마크를 구축했습니다. 실험 결과, 현재의 MLLM은 테이블 구조를 안정적으로 재구성할 수 있지만, 정확한 값 복구에는 어려움을 겪는 것으로 나타났습니다. 이에 우리는 인간 중심적인 관점에서 차트 데이터 추출을 재검토하고, 추출 과정이 사람들이 차트를 읽는 방식과 유사한 점진적인 학습 과정을 따라야 한다고 주장합니다. 제안하는 학습 프레임워크는 수치 정확도를 크게 향상시켜 70억 개의 파라미터를 가진 모델로 최첨단 성능을 달성했습니다. 사용자 연구 결과, 개발된 모델은 신뢰할 수 있는 차트 데이터 추출을 위한 혼합형 워크플로우를 효과적으로 지원하는 것으로 나타났습니다.
Chart data extraction, which reverse-engineers data tables from chart images, is essential for reproducibility, analysis, retrieval, and redesign. Existing interactive tools are reliable but tedious, and mixed-initiative systems, while more efficient, lack generalizability. Recent multimodal large language models (MLLMs) offer a unified interface for chart interpretation, yet their ability to extract accurate data tables, especially without visible labels, remains unclear. We build a benchmark featuring diverse real-world charts without data labels to evaluate this capability. Results show that, while current MLLMs reliably reconstruct table structures, they struggle with precise value recovery. To address this, we revisit chart data extraction from a human-centered perspective and argue that extraction should follow a progressive learning process similar to how people read charts. Our training framework substantially improves numerical accuracy, achieving state-of-the-art performance with a 7B-parameter model. A user study further shows that our model effectively supports mixed-initiative workflows for reliable chart data extraction.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.