2608.04504v1 Aug 05, 2026 cs.CV

GeoReward: 비전-언어 모델의 교차 시장 선호도 예측에서 맥락 변수 과대 추정 완화

GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction

Xiaoyi Zeng
Xiaoyi Zeng
Citations: 137
h-index: 3
Weiru Zhang
Weiru Zhang
Citations: 43
h-index: 3
Hui Cai
Hui Cai
Citations: 0
h-index: 0
Shuo Liu
Shuo Liu
Citations: 0
h-index: 0

비전-언어 모델은 다양한 멀티모달 작업에서 뛰어난 성능을 보이지만, 미묘하지만 중요한 문제점인 맥락 변수의 과대 추정 경향이 있습니다. 이러한 문제는 특히 서로 다른 지리적 시장에서의 광고 이미지 선호도를 예측하는 등 실제 응용 분야에서 두드러집니다. 예를 들어, 비전-언어 모델이 여러 국가를 위해 맞춤 제작된 두 가지 제품 이미지 중에서 선택하도록 요청받으면, 종종 일관된 결과를 내놓으며 실제 지역별 차이를 무시합니다. 이는 제품 속성이나 밀집된 이미지 영역과 같은 광범위한 신호가 시장별 맥락을 담고 있는 소수의 중요한 토큰을 압도하기 때문에 발생합니다. 이러한 맥락 변수 과대 추정(CVE) 문제를 해결하기 위해, 우리는 여러 국가에서 수집된 실제 광고 콘텐츠와 클릭률 데이터를 활용하여 새로운 멀티모달 데이터 세트를 구축했습니다. 그리고 다양한 지리적 시장에서의 광고 이미지 선호도를 예측하도록 설계된 보상 모델인 GeoReward를 제안합니다. GeoReward는 다음 세 가지 주요 메커니즘을 통합합니다: (1) 시장 인지 검색 증강(Market-Aware Retrieval Augmentation), (2) 맥락 기반 시각적 조절(Context-Guided Visual Modulation), (3) 선택적 민감도 손실(Selective Sensitivity Loss). 또한, GeoReward가 비전-언어 모델의 강화 학습을 통해 텍스트-이미지 모델을 위한 배경 디자인 생성을 안내하여 시장에 적합한 광고 콘텐츠를 생성하는 방법을 보여줍니다. 실험 결과는 우리의 프레임워크가 CVE를 완화하고 기존 방법보다 우수한 성능을 보임을 입증합니다. 본 연구는 비전-언어 모델의 지배적인 시각적 특징에 대한 체계적인 편향을 진단할 뿐만 아니라, 희소한 맥락 변수가 의사 결정에 중요한 영향을 미치는 응용 분야를 위한 맞춤형 솔루션을 제공합니다.

Original Abstract

Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables. This issue, which we term Contextual Variable Overestimation (CVE), becomes particularly evident in real-world applications such as predicting advertisement image preferences across diverse geographic markets. For instance, when a VLM is asked to choose between two product images tailored for different countries, it often defaults to a consistent output, ignoring ground-truth regional variations. This collapse occurs because pervasive high-volume signals, such as product attributes and dense image patches, overwhelm the few but critical tokens that encode market-specific context. To address CVE, we first collect a new multimodal dataset of real advertising creatives and their click-through performance across multiple countries. We then introduce GeoReward, a reward model designed to predict ad image preferences across diverse geographic markets. GeoReward integrates three purpose-built mechanisms: (1) Market-Aware Retrieval Augmentation, (2) Context-Guided Visual Modulation, (3) Selective Sensitivity Loss. Furthermore, we demonstrate how GeoReward can guide the fine-tuning of RL for a VLM to generate background designs for text-to-image models, producing market-aware advertising creatives. Experiments validate that our framework mitigates CVE and outperforms existing baselines. This work not only diagnoses a systematic bias in VLMs toward dominant perceptual features but also delivers a targeted solution for applications where sparse contextual variables govern decision-making.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!