동적이고 다중 턴 상호작용을 통한 시각-언어 모델의 맥락화된 평가
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
다중 모드 대규모 언어 모델(MLLM)은 기존 벤치마크에서 상당한 발전을 이루었지만, 실제 환경에서의 효과성은 여전히 불확실합니다. 이러한 격차는 통제되고 정적인 환경에서의 벤치마크와 실제 응용 분야의 역동적이고 상호 작용적이며 맥락화된 특성 간의 근본적인 불일치에서 비롯됩니다. 이 격차를 해소하기 위해, 우리는 CEDI(Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions)라는 프레임워크를 제안합니다. CEDI는 평가를 피 평가 모델, 자동화된 검사자 및 평가자로 구성된 3자 상호작용으로 재구성합니다. 검사자는 그래프 기반 표현을 통해 안내되는 다중 턴의 반구조적 대화를 수행하며, 상태 공간 전환을 통해 다양한 전략(예: 명확화 요청 또는 적대적 탐색)을 사용하여 성능 증거를 이끌어냅니다. 우리는 CEDI를 시각적 환각 평가에 적용했습니다. 여러 모델, 다양한 설정, 데이터 세트 및 도메인에서의 실증적인 결과는 맥락화되고 상호 작용적인 평가가 기존의 정적 평가보다 훨씬 더 많은 수의 환각을 드러낼 뿐만 아니라, 실제 사용 사례에서 발생하는 환각과 더욱 유사한 환각을 발견한다는 것을 보여줍니다. 또한, 환각은 종종 긴 문맥 속에서, 그리고 자기 강화되는 대화 기록을 통해 축적되며, 전제 부정 또는 거부를 요구하는 질문에 대해 모델이 특히 취약하다는 것을 확인했습니다. 이러한 결과들은 CEDI가 MLLM의 기능에 대한 현실적이고 체계적이며 생태적으로 타당한 평가를 향한 중요한 단계임을 강조합니다. 코드: github.com/williamium3000/cedi.
Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions), a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured conversation guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies, from clarification requests to adversarial probes, to elicit performance evidence. We apply CEDI to visual hallucinations. Empirical results across multiple models, diverse settings, datasets, and domains show that contextualized, interactive evaluations reveal not only significantly more hallucinations than conventional static evaluation but also ones that more closely resemble those arising in practical use cases. We further show that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal. Together, these findings highlight CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs' capabilities. Code is available at github.com/williamium3000/cedi.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.