2607.25911v1 Jul 28, 2026 cs.HC

AnnoBench: 시각화 주석 생성 성능 평가 도구

AnnoBench: A Benchmark for Visualization Annotation Generation

Md Rahat-uz-Zaman
Md Rahat-uz-Zaman
Citations: 1
h-index: 1
Md Dilshadur Rahman
Md Dilshadur Rahman
Citations: 39
h-index: 4
Andrew McNutt
Andrew McNutt
Citations: 2
h-index: 1
Paul Rosen
Paul Rosen
Citations: 79
h-index: 4

시각화 주석 생성은 시각적, 의미론적, 스타일적 제약을 동시에 만족시켜야 하는 까다로운 작업입니다. 이러한 조건 중 하나라도 충족되지 않으면 주석의 유용성이 크게 저하되어 가독성을 떨어뜨리거나 부정확하게 만들고, 심지어 시각적으로 어울리지 않게 됩니다. 현재 다양한 시각화 도구 및 자동화 기술이 개발되고 있지만, 대부분의 연구는 특정 범위에 집중하거나 주석 생성 자체를 주요 연구 대상으로 하지 않아 이러한 조건들이 제대로 충족되는지를 평가하는 벤치마크나 평가 프레임워크가 부족합니다. 본 논문에서는 AnnoBench라는 시각화 주석 생성을 위한 벤치마크를 소개합니다. AnnoBench는 이 분야의 내재된 어려움을 체계적이고 검증 가능한 방식으로 구현하여 제공합니다. An노Bench는 전문 데이터 저널리즘 및 시각화 갤러리에서 가져온 시각화 자료와 주석 생성 작업을 결합하며, 네 가지 표현 형식, 다섯 가지 차트 설명 조건, 그리고 두 가지 프롬프트 명세 수준을 포함합니다. 본 벤치마크는 VLM-as-a-judge 시스템을 사용하여, 사람의 평가와 일관성을 갖도록 조정된 모델을 통해 실행됩니다. 우리는 입력 표현 방식, 의미론적 맥락, 프롬프트 구체성 및 모델 선택이 주석 품질에 미치는 영향을 분석하기 위해 네 가지 요인을 하나씩 변경하는 실험을 수행했습니다. 본 연구는 시각화 주석 자동화, 도구 개발 및 시각화 생성 파이프라인 발전에 필요한 기반을 제공합니다.

Original Abstract

Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints. Failure to meet any of these conditions severely undermines the utility of an annotation, rendering it challenging to read, inaccurate, or visually discordant. Despite a growing body of annotation tools and automations, no existing benchmark or evaluation framework tests whether these conditions are met because of their scope and annotation not being the focus of their studies. We introduce AnnoBench, a benchmark for visualization annotation that materializes the inherent challenges of this domain in a structured and testable manner. AnnoBench pairs visualizations from professional data journalism and visualization galleries with annotation tasks, spanning four representation formats, five chart description conditions, and two prompt specification levels. The benchmark is executed via VLM-as-a-judge, using models aligned with manual human assessment. We evaluate the benchmark via four one-factor-at-a-time experiments, exploring the effects of input representation, semantic context, and prompt specificity, and model selection on annotation quality. This work provides a foundation for advancing annotation automation, tooling, and visualization-generation pipelines.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!