ENTRAP-VL: 시각-언어 모델의 이중 맥락적 영향력 분석 도구
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
맥락적 영향력은 모델이 입력에 포함된 보조 정보를 바탕으로 출력을 생성하는 경향을 의미하며, 해당 정보가 관련성이 있거나, 진실되거나, 심지어 의미가 있는지 여부에 관계없이 발생합니다. 최근 연구에서 단일 모드 언어 모델에서 이러한 현상이 확인되고 메커니즘적으로 설명되었습니다. 반면, 시각-언어 모델(VLM)에서 맥락적 영향력이 어떻게 나타나는지는 거의 연구되지 않았으며, 이를 조사하기 위한 특수 목적의 도구가 부족합니다. 본 연구에서는 VLM에서의 맥락적 영향력 분석이 기존 텍스트 기반 벤치마크를 다중 모드 환경으로 단순히 확장하는 것 이상이라는 입장을 취합니다. 이는 분류학적으로 구조화되고 이중 모드 방식으로 설계된 도구를 필요로 하며, 이 도구의 조건은 해당 항목(텍스트 스트림에서 묘사된 이미지 또는 시각 스트림에서 제시된 텍스트 질의)을 중심으로 구성되어야 합니다. 우리는 VLM으로의 전환이 점진적인 개선이 아닌 실질적인 변화라고 주장합니다. 이는 맥락적 영향력을 텍스트와 시각, 두 가지 독립적인 요인에 의해 유발될 수 있는 이중 현상으로 만들고, 이전 연구에서 다루지 않았던 '세계적으로 가능하지만 묘사된 장면에는 거짓인 정보'라는 새로운 차원을 제시합니다. 이러한 주장을 구체화하고 실질적인 활용을 가능하게 하기 위해, 우리는 시각-언어 맥락적 영향력 평가 도구인 ENTRAP-VL(ENTRainment Assessment Probe for Vision and Language)을 개발했습니다. 이는 8가지 범주에 걸쳐 수동으로 구성된 1,500개의 항목 데이터 세트로, 두 가지 축(항목과의 맥락 연관성 및 진실 여부와의 관계)으로 분류된 분류 체계를 따르며, 텍스트 기반 영향력 평가 스트림(8가지 조건)과 시각 기반 영향력 평가 스트림(3가지 조건)으로 구성되어 있습니다. 본 연구는 특정 모델의 맥락적 영향력을 측정하는 것을 목표로 하지 않으며, 분석 도구 자체, 이를 뒷받침하는 분류 체계, 그리고 이 도구를 통해 가능해지는 평가 프로토콜을 제공하여 연구 커뮤니티가 해당 현상을 엄밀하게 조사할 수 있도록 합니다. 개발된 데이터 세트와 관련 문서는 공개적으로 배포될 예정입니다.
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.