시각적 설득: 비전-언어 모델의 의사 결정에 어떤 요인이 영향을 미치는가?
Visual Persuasion: What Influences Decisions of Vision-Language Models?
인터넷에는 수많은 이미지가 존재하며, 과거에는 인간을 위해 만들어졌지만 현재는 비전-언어 모델(VLMs)을 사용하는 에이전트들이 점점 더 많이 해석하고 있습니다. 이러한 에이전트들은 클릭, 추천 또는 구매와 같은 시각적 의사 결정을 대규모로 수행합니다. 그러나 그들의 시각적 선호도에 대한 정보는 아직 부족합니다. 본 연구에서는 VLMs를 제어된 이미지 기반 선택 과제에 배치하고 입력값을 체계적으로 변경함으로써 이를 연구하기 위한 프레임워크를 제시합니다. 핵심 아이디어는 에이전트의 의사 결정 함수를 잠재적인 시각적 효용성으로 간주하고, 체계적으로 편집된 이미지 간의 선택을 통해 이를 추론하는 것입니다. 일반적인 이미지, 예를 들어 제품 사진을 기반으로, 시각적 프롬프트 최적화 방법을 제안합니다. 이 방법은 텍스트 최적화 방법을 활용하여 이미지 생성 모델(예: 합성, 조명 또는 배경)을 사용하여 시각적으로 타당한 수정 사항을 반복적으로 제안하고 적용합니다. 그런 다음, 어떤 수정 사항이 선택 확률을 증가시키는지 평가합니다. 최첨단 VLMs에 대한 대규모 실험을 통해, 최적화된 수정 사항이 직접 비교에서 선택 확률을 크게 변화시킨다는 것을 입증했습니다. 이러한 선호도를 설명하기 위해 자동 해석 파이프라인을 개발하여 선택을 유발하는 일관된 시각적 특징을 식별했습니다. 본 연구는 이미지 기반 AI 에이전트의 잠재적인 취약점을 실질적이고 효율적으로 파악하는 방법을 제공하며, 이는 야생 환경에서 암묵적으로 발견될 수 있는 안전 문제를 사전에 예방하고, 보다 적극적인 감사 및 관리를 지원할 수 있습니다.
The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (VLMs). These agents make visual decisions at scale, deciding what to click, recommend, or buy. Yet, we know little about the structure of their visual preferences. We introduce a framework for studying this by placing VLMs in controlled image-based choice tasks and systematically perturbing their inputs. Our key idea is to treat the agent's decision function as a latent visual utility that can be inferred through revealed preference: choices between systematically edited images. Starting from common images, such as product photos, we propose methods for visual prompt optimization, adapting text optimization methods to iteratively propose and apply visually plausible modifications using an image generation model (such as in composition, lighting, or background). We then evaluate which edits increase selection probability. Through large-scale experiments on frontier VLMs, we demonstrate that optimized edits significantly shift choice probabilities in head-to-head comparisons. We develop an automatic interpretability pipeline to explain these preferences, identifying consistent visual themes that drive selection. We argue that this approach offers a practical and efficient way to surface visual vulnerabilities, safety concerns that might otherwise be discovered implicitly in the wild, supporting more proactive auditing and governance of image-based AI agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.