시각-언어 모델에서의 정보 출처 모니터링
Source-Modality Monitoring in Vision-Language Models
본 논문에서는 다중모드 모델이 입력 정보의 출처를 추적하고 전달하는 능력인 '정보 출처 모니터링'을 정의하고 연구합니다. 우리는 정보 출처 모니터링을 더 일반적인 '바인딩 문제'의 한 예시로 간주하며, 모델이 사용자가 제공한 프롬프트에서 '이미지'와 같은 단어를 입력 및 컨텍스트의 특정 구성 요소(즉, 실제 이미지)에 연결할 때 구문 정보와 의미 정보를 얼마나 활용하는지를 평가합니다. 11개의 시각-언어 모델(VLM)을 대상으로 목표 모드 정보 검색 작업을 수행한 실험에서, 구문 정보와 의미 정보 모두 중요한 역할을 한다는 것을 확인했지만, 모달리티의 분포가 매우 다를 경우 의미 정보가 더 큰 영향을 미치는 경향이 있음을 발견했습니다. 본 연구 결과는 모델의 안정성 및 점점 더 복잡해지는 다중모드 에이전트 시스템 개발에 대한 함의를 논합니다.
We define and investigate source-modality monitoring -- the ability of multimodal models to track and communicate the input source from which pieces of information originate. We consider source-modality monitoring as an instance of the more general binding problem, and evaluate the extent to which models exploit syntactic vs. semantic signals in order to bind words like image in a user-provided prompt to specific components of their input and context (i.e., actual images). Across experiments spanning 11 vision-language models (VLMs) performing target-modality information retrieval tasks, we find that both syntactic and semantic signals play an important role, but that the latter tend to outweigh the former in cases when modalities are highly distinct distributionally. We discuss the implications of these findings for model robustness, and in the context of increasingly multimodal agentic systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.