2607.04727v1 Jul 06, 2026 cs.SE

Dashboard2Code: 대화형 대시보드 재구성을 통한 다중 모달 모델 평가

Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards

Qingfu Zhu
Qingfu Zhu
Citations: 553
h-index: 11
Shiqi Zhou
Shiqi Zhou
Citations: 2
h-index: 1
Wanxiang Che
Wanxiang Che
Citations: 357
h-index: 10
Qiguang Chen
Qiguang Chen
SCIR
Citations: 1,830
h-index: 22
Ziyu Han
Ziyu Han
Citations: 1
h-index: 1
Tianhao Niu
Tianhao Niu
Citations: 9
h-index: 2
Hengjie Fang
Hengjie Fang
Citations: 0
h-index: 0
B. Shan
B. Shan
Citations: 0
h-index: 0

최근 다중 모달 대규모 언어 모델의 발전으로 자동 데이터 시각화 생성 기술이 빠르게 발전하고 있지만, 기존 연구는 주로 정적 차트에 집중하며 실제 데이터 탐색에 널리 사용되는 대화형 대시보드를 간과하는 경향이 있습니다. 본 논문에서는 모델이 능동적으로 대화형 대시보드를 탐색하고, 자체적인 상호 작용(예: 클릭 및 필터링)으로부터 피드백을 수집하여 통합하며, 목표 대시보드를 재현하는 코드를 생성하도록 요구하는 새로운 작업인 Dashboard2Code를 소개합니다. 포괄적인 평가를 지원하기 위해, 본 논문에서는 Dashboard2Code를 위한 Plotly+Dash 벤치마크인 DashboardMimic을 제시합니다. DashboardMimic은 세 가지 난이도 수준으로 구성된 180개의 신중하게 설계되고 수동으로 검증된 대시보드-코드 쌍과 함께, 8가지 일반적인 실제 상호 작용 패턴을 포함하고 있습니다. 또한, 본 논문에서는 시각적 일관성과 상호 작용 일관성을 평가하기 위해 코드 의미 분석과 동적 상호 작용 기반 테스트를 결합한 자동화된 평가 프레임워크를 제안하며, 이 프레임워크는 인간의 판단과 높은 상관 관계를 보입니다. 다양한 공개 및 비공개 다중 모달 모델에 대한 실험 결과, 가장 강력한 시스템조차도 복잡도가 높은 대시보드에서 어려움을 겪으며, Dashboard2Code 작업에서 오픈 소스 모델과 비공개 모델 간에는 상당한 성능 격차가 존재하는 것으로 나타났습니다.

Original Abstract

Automatic data visualization generation has advanced rapidly with multi-modal large language models, yet existing efforts largely focus on static charts and overlook the interactive dashboards commonly used for real-world data exploration. We introduce Dashboard2Code, a novel task that requires a model to proactively explore an interactive dashboard, acquire and integrate feedback from its own interactions (e.g., clicking and filtering), and generate code that reproduces the target dashboard. To support comprehensive evaluation, we present DashboardMimic, the first Plotly+Dash benchmark for Dashboard2Code, comprising 180 carefully designed and manually verified dashboard-code pairs spanning three difficulty levels and covering eight common real-world interaction patterns. We further propose an automated evaluation framework tailored to dashboards that combines code semantic analysis with dynamic interaction-based testing to assess visual and interaction consistency, showing strong agreement with human judgments. Experiments across a range of open- and closed-source multi-modal models reveal that even the strongest systems struggle on high-complexity dashboards and that a substantial performance gap remains between open-source and closed-source models on the Dashboard2Code task.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!