2608.03464v1 Aug 04, 2026 cs.AI

ChartAnno: 차트 어노테이션 생성을 위한 멀티모달 대규모 언어 모델 평가

ChartAnno: Evaluating MLLMs for Chart Annotation Generation

Zekai Shao
Zekai Shao
Citations: 692
h-index: 9
Xingchen Zeng
Xingchen Zeng
Hong Kong University of Science and Technology (Guangzhou)
Citations: 94
h-index: 4
Xiaoliang Fu
Xiaoliang Fu
Fudan University
Citations: 48
h-index: 4
Ziyue Lin
Ziyue Lin
Citations: 65
h-index: 5
Yue Guo
Yue Guo
Citations: 0
h-index: 0
Lidan Tan
Lidan Tan
Citations: 0
h-index: 0
Xin Lin
Xin Lin
Citations: 0
h-index: 0
Yi Shan
Yi Shan
Citations: 35
h-index: 3
Bongshin Lee
Bongshin Lee
Citations: 140
h-index: 4
Zhenghan Chen
Zhenghan Chen
Citations: 0
h-index: 0
Xinyuan Liu
Xinyuan Liu
Citations: 0
h-index: 0
Fen Wang
Fen Wang
Citations: 15
h-index: 1
Siming Chen
Siming Chen
Citations: 61
h-index: 4

멀티모달 대규모 언어 모델(MLLM)은 차트 이해, 생성 및 편집 분야에서 상당한 발전을 이루었지만, 기존 차트에 대한 어노테이션을 수행하는 능력은 아직 충분히 연구되지 않았습니다. 차트 어노테이션은 흔히 사용되는 의사소통 작업이지만, 모델이 의도된 메시지를 추론하고, 차트의 의미를 해석하며, 적절한 텍스트 또는 그래픽 요소를 배치해야 하기 때문에 어려운 과제입니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 MLLM의 차트 어노테이션 생성 성능을 평가하는 데 사용될 수 있는 벤치마크인 ChartAnno를 소개합니다. ChartAnno는 세 가지 수준의 지시 명확성을 가진 1,200개의 실제 차트와 해당 코드 및 어노테이션 지침으로 구성되어 있습니다. 우리는 10개의 대표적인 MLLM을 두 가지 주요 입력 방식으로 평가했습니다: (1) 차트 코드만 사용하고, (2) 차트 코드와 차트 이미지를 모두 사용하는 방식입니다. 또한, 차트 이미지만을 사용한 실험도 포함하여 분석의 폭을 넓혔습니다. 결과는 자체 모델이 여전히 전반적으로 더 우수하지만, 대규모 오픈 소스 모델들이 그 격차를 줄이고 있음을 보여줍니다. 더욱 구체적인 지시는 어노테이션 품질을 향상시키는 데 도움이 되지만, 추상적인 의도를 파악하는 것은 현재의 MLLM에게 여전히 가장 어려운 과제입니다. 차트 이미지를 제공하면 전반적으로 제한적인 개선 효과가 나타나며, 주로 디자인 관련 지표에서 개선이 관찰됩니다. 이러한 결과는 차트 어노테이션 생성이 의미론적 이해와 효과적인 어노테이션 설계가 필요한 어려운 작업임을 강조합니다. 코드 및 데이터는 향후 버전에 공개될 예정입니다.

Original Abstract

Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative task, requiring models to infer intended messages, interpret chart semantics, and place appropriate textual or graphical elements. To address this gap, we introduce ChartAnno, a benchmark for evaluating MLLMs on chart annotation generation. It contains 1,200 real-world charts with paired code and annotation instructions across three levels of instruction specificity. We evaluate 10 representative MLLMs under two primary input settings: (1) chart code alone and (2) both chart code and chart image, and further include a chart image-only ablation study. Results show that proprietary models remain stronger overall, although large-scale open-source models narrow the gap. More specific instructions improve annotation quality, while inferring abstract intent remains most difficult for current MLLMs. Providing chart images brings limited overall gains, with improvements mainly appearing in design-related metrics. These findings highlight chart annotation generation as a challenging task requiring semantic grounding and effective annotation design. Code and data will be released in a future version.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!