2604.22038v2 Apr 23, 2026 cs.CL

시각-언어 모델에서의 정보 출처 모니터링

Source-Modality Monitoring in Vision-Language Models

Ellie Pavlick
Ellie Pavlick
Citations: 3,877
h-index: 11
Etha Tianze Hua
Etha Tianze Hua
Brown University
Citations: 35
h-index: 2
Tian Yun
Tian Yun
Brown University
Citations: 3,106
h-index: 9

본 논문에서는 다중모드 모델이 입력 정보의 출처를 추적하고 전달하는 능력인 '정보 출처 모니터링'을 정의하고 연구합니다. 우리는 정보 출처 모니터링을 더 일반적인 '바인딩 문제'의 한 예시로 간주하며, 모델이 사용자가 제공한 프롬프트에서 '이미지'와 같은 단어를 입력 및 컨텍스트의 특정 구성 요소(즉, 실제 이미지)에 연결할 때 구문 정보와 의미 정보를 얼마나 활용하는지를 평가합니다. 11개의 시각-언어 모델(VLM)을 대상으로 목표 모드 정보 검색 작업을 수행한 실험에서, 구문 정보와 의미 정보 모두 중요한 역할을 한다는 것을 확인했지만, 모달리티의 분포가 매우 다를 경우 의미 정보가 더 큰 영향을 미치는 경향이 있음을 발견했습니다. 본 연구 결과는 모델의 안정성 및 점점 더 복잡해지는 다중모드 에이전트 시스템 개발에 대한 함의를 논합니다.

Original Abstract

We define and investigate source-modality monitoring -- the ability of multimodal models to track and communicate the input source from which pieces of information originate. We consider source-modality monitoring as an instance of the more general binding problem, and evaluate the extent to which models exploit syntactic vs. semantic signals in order to bind words like image in a user-provided prompt to specific components of their input and context (i.e., actual images). Across experiments spanning 11 vision-language models (VLMs) performing target-modality information retrieval tasks, we find that both syntactic and semantic signals play an important role, but that the latter tend to outweigh the former in cases when modalities are highly distinct distributionally. We discuss the implications of these findings for model robustness, and in the context of increasingly multimodal agentic systems.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!