차트 지원 또는 모델 제공? 접근 가능한 시각화를 위한 MLLM 생성 주장의 검토
Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization
다중 모드 대규모 언어 모델(MLLM)은 시각화 패턴을 외부 원인, 결과 및 도메인 지식과 연결할 수 있지만, 이러한 해석의 증거 기반은 종종 불분명합니다. 본 연구는 4가지 출처에서 추출한 102개의 시각화 자료와 3개의 MLLM 모델, 그리고 이미지 접근성, 출처별 접근 가능한 차트 정보 제공 여부, 그리고 숨겨진 정보 프레임의 변화를 반영하는 4가지 입력 조건에 대한 탐색적 연구입니다. 총 1,224개의 설명에 대해, 모델이 부여한 직접(DIRECT), 유도(DERIVED) 및 추론(SPECULATIVE) 레이블을 분석하고 숫자적 일관성을 자동 검증했습니다. 접근 가능한 차트 정보는 Gemini와 GPT 모델의 경우 직접적인 주장을 증가시키고 일부 모델의 숫자적 일관성을 향상시켰습니다. 전체 컨텍스트에 이미지를 추가하는 것이 일관된 숫자적 이점을 제공하지 않았으며, 숨겨진 컨텍스트 프롬프트가 항상 신중한 표현을 증가시키지는 않았습니다. 프롬프트를 통해 정의된 실제 중요성(Real-World Significance) 섹션은 여전히 주로 추론적인 내용으로 구성되었습니다. 이러한 결과는 제공된 증거에 의해 뒷받침되는 주장을 모델이 제공하는 해석과 구별하는 접근 가능한 설명 시스템 개발의 필요성을 강조합니다.
Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that vary access to the image, source-specific accessible chart context, and withheld-context framing. Across 1,224 descriptions, we analyze model-attributed DIRECT, DERIVED, and SPECULATIVE labels and conduct an automated audit of numeric agreement. Accessible chart context shifted Gemini and GPT toward DIRECT claims and improved numeric agreement for some models. Adding the image to the full context did not yield a consistent numeric benefit, and the withheld-context prompt did not reliably increase cautious language. The prompt-defined Real-World Significance section remained predominantly SPECULATIVE. These results motivate accessible description systems that distinguish claims supported by supplied evidence from model-supplied interpretation
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.