SPOT-E: 시각적 스포트라이트를 활용한 테스트 단계에서의 엔트로피 제어 - 동결된 멀티모달 모델을 위한 방법
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
비전-언어 모델(VLM)은 종종 중요한 증거가 작고, 국소화되어 있으며, 간과하기 쉬운 경우와 같이, 결정적인 시각적 증거를 필요로 하는 작업에서 성능이 저조합니다. 이러한 문제는 고차원 추론 능력이 유지되더라도 증거 판독 실패로 이어질 수 있습니다. 기존의 추론 시간 시각적 개입은 재학습 없이 근거 설정(grounding)을 향상시킬 수 있지만, 대부분 폐루프 방식이며 강조된 증거가 실제로 사용되는지 확인하는 메커니즘이 부족합니다. 본 연구에서는 모델 내부 피드백 신호인 답변 범위 예측 엔트로피를 분석하고, 단순한 엔트로피 최소화는 모호할 수 있음을 보여줍니다. 왜냐하면 낮은 엔트로피는 증거 기반의 확신 또는 단축 경로(shortcut)로 인해 발생할 수 있기 때문입니다. 이러한 모호성을 해결하기 위해, 우리는 낮은 엔트로피 지점(anchor)과 답변 불확실성을 줄이면서 기본적으로 높은 신뢰도를 가진 토큰을 유지하는 엔트로피 형성 목표 함수를 도입했습니다. 우리는 이 원리를 SPOT-E라는 플러그 앤 플레이 방식의 테스트 시간 기법에 구현했습니다. SPOT-E는 각 인스턴스별로 경량화된 그룹 상대 정책 최적화(GRPO) 기반 튜닝을 통해 질문에 따라 생성되는 시각적 스포트라이트를 생성합니다. 다양한 벤치마크와 VLM 패밀리에서 SPOT-E는 일관된 성능 향상과 시각적 왜곡에 대한 개선된 안정성을 보여줍니다. 코드 공개 위치: https://github.com/YinBo0927/SPOT-E
Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evidence readout even when high-level reasoning is intact. Prior inference-time visual interventions can improve grounding without retraining, but they are largely open-loop and lack a mechanism to verify whether highlighted evidence is actually used. We study answer-span prediction entropy as a model-internal feedback signal and show that naive entropy minimization is ambiguous, since low entropy may arise from evidence-grounded confidence or shortcut collapse. To resolve this ambiguity, we introduce low-entropy anchors and an entropy-shaping objective that reduces answer uncertainty while preserving baseline high-confidence tokens. We instantiate this principle in SPOT-E, a plug-and-play test-time method that produces question-conditioned spotlights, optimized per instance via light-weight tuning based on Group Relative Policy Optimization (GRPO). Across all benchmarks and different VLM families, SPOT-E yields consistent gains and improved robustness under visual corruptions. Code is publicly available at: \url{https://github.com/YinBo0927/SPOT-E}
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.