텍스트 기반 편향 제거: 눈으로 보고 판단하세요 - 시각적 반상식 추론을 위한 텍스트 연계 교차 모달 전이
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
멀티모달 대규모 언어 모델(MLLM)의 시각적 추론 능력은 다운스트림 애플리케이션, 특히 일반적인 가정에서 벗어난 반상식 추론에 매우 중요합니다. 최근 연구에서는 주로 시각적 입력 개선을 통해 시각적 반상식 추론 능력을 향상시켜 왔는데, 이는 모델의 실패 원인이 충분한 시각적 정보 부족이라는 가정에 기반합니다. 그러나 우리의 실증 분석 결과, 문제의 병목 지점은 시각적 인식 자체가 아닙니다. MLLM은 이미 관련 시각적 증거를 파악하고 있으며, 정답은 모델의 디코딩 공간 내에 존재합니다. 대신, 공유 언어 디코더는 특히 낮은 빈도의 사실 기반 시나리오에서 우세한 언어적 선입견을 따르면서 선입견과 증거 간의 충돌을 해결합니다. 이러한 점에 착안하여, 우리는 먼저 텍스트 기반 데이터 구축 파이프라인을 제안합니다. 이 파이프라인의 핵심 구성 요소인 Fact-Frequency Distillation (FFD)은 상식 사실의 선입견 강도를 추정하고 검증된 반상식 시나리오를 고품질 텍스트 코퍼스로 변환합니다. 이렇게 구축된 코퍼스를 기반으로, 우리는 TACT라는 텍스트 연계 후속 학습 프레임워크를 소개합니다. TACT는 시각적 학습 데이터 없이 공유 언어 디코더의 편향을 제거합니다. TACT는 증거 추론과 선입견 기반 추론 경로를 서로 다른 최적화 단계로 분리하여, 디코더가 선입견과 증거 간의 충돌을 해결할 수 있도록 합니다. 다양한 시각적 반상식 벤치마크에서 TACT는 시각적 추론 능력을 크게 향상시키면서도 일반적인 기능을 유지하며, 효과적인 텍스트-시각 교차 모달 전이를 보여줍니다.
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models to reason beyond common assumptions. Recent studies mainly improve visual counter-commonsense reasoning by enhancing visual inputs, following the assumption that failures originate from insufficient visual grounding. However, our empirical analysis reveals that the bottleneck is not visual perception. MLLMs already capture the relevant visual evidence, and the correct answer exists in their decoding space. Instead, the shared language decoder resolves prior--evidence conflicts by favoring dominant language priors, especially for low-frequency factual scenarios. Motivated by this, we first propose a text-anchored data construction pipeline, whose core component, Fact-Frequency Distillation (FFD), estimates the prior strength of commonsense facts and distills verified counter-commonsense scenarios into a high-quality text corpus. Building upon this corpus, we introduce TACT, a text-anchored post-training framework that debiases the shared language decoder without requiring any visual training data. TACT routes evidence-following and prior-driven reasoning trajectories into different optimization stages, enabling the decoder to resolve prior--evidence conflicts. Across counter-commonsense visual benchmarks, TACT substantially improves visual reasoning while preserving general capabilities, demonstrating effective text-to-vision cross-modal transfer.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.