2606.20244v1 Jun 18, 2026 cs.CV

SPOT-E: 시각적 스포트라이트를 활용한 테스트 단계에서의 엔트로피 제어 - 동결된 멀티모달 모델을 위한 방법

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

Bo Yin
Bo Yin
Citations: 11
h-index: 2
Chengming Xu
Chengming Xu
Citations: 228
h-index: 8
Shuicheng Yan
Shuicheng Yan
Citations: 113
h-index: 6
Xiaobin Hu
Xiaobin Hu
Citations: 32
h-index: 3
Jiang-She Zhang
Jiang-She Zhang
Citations: 179
h-index: 8
Cheng Tan
Cheng Tan
Citations: 10
h-index: 2
Peng-Tao Jiang
Peng-Tao Jiang
Citations: 35
h-index: 3
Ruolin Shen
Ruolin Shen
Citations: 3
h-index: 1
Mo Yang
Mo Yang
Citations: 18
h-index: 2

비전-언어 모델(VLM)은 종종 중요한 증거가 작고, 국소화되어 있으며, 간과하기 쉬운 경우와 같이, 결정적인 시각적 증거를 필요로 하는 작업에서 성능이 저조합니다. 이러한 문제는 고차원 추론 능력이 유지되더라도 증거 판독 실패로 이어질 수 있습니다. 기존의 추론 시간 시각적 개입은 재학습 없이 근거 설정(grounding)을 향상시킬 수 있지만, 대부분 폐루프 방식이며 강조된 증거가 실제로 사용되는지 확인하는 메커니즘이 부족합니다. 본 연구에서는 모델 내부 피드백 신호인 답변 범위 예측 엔트로피를 분석하고, 단순한 엔트로피 최소화는 모호할 수 있음을 보여줍니다. 왜냐하면 낮은 엔트로피는 증거 기반의 확신 또는 단축 경로(shortcut)로 인해 발생할 수 있기 때문입니다. 이러한 모호성을 해결하기 위해, 우리는 낮은 엔트로피 지점(anchor)과 답변 불확실성을 줄이면서 기본적으로 높은 신뢰도를 가진 토큰을 유지하는 엔트로피 형성 목표 함수를 도입했습니다. 우리는 이 원리를 SPOT-E라는 플러그 앤 플레이 방식의 테스트 시간 기법에 구현했습니다. SPOT-E는 각 인스턴스별로 경량화된 그룹 상대 정책 최적화(GRPO) 기반 튜닝을 통해 질문에 따라 생성되는 시각적 스포트라이트를 생성합니다. 다양한 벤치마크와 VLM 패밀리에서 SPOT-E는 일관된 성능 향상과 시각적 왜곡에 대한 개선된 안정성을 보여줍니다. 코드 공개 위치: https://github.com/YinBo0927/SPOT-E

Original Abstract

Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evidence readout even when high-level reasoning is intact. Prior inference-time visual interventions can improve grounding without retraining, but they are largely open-loop and lack a mechanism to verify whether highlighted evidence is actually used. We study answer-span prediction entropy as a model-internal feedback signal and show that naive entropy minimization is ambiguous, since low entropy may arise from evidence-grounded confidence or shortcut collapse. To resolve this ambiguity, we introduce low-entropy anchors and an entropy-shaping objective that reduces answer uncertainty while preserving baseline high-confidence tokens. We instantiate this principle in SPOT-E, a plug-and-play test-time method that produces question-conditioned spotlights, optimized per instance via light-weight tuning based on Group Relative Policy Optimization (GRPO). Across all benchmarks and different VLM families, SPOT-E yields consistent gains and improved robustness under visual corruptions. Code is publicly available at: \url{https://github.com/YinBo0927/SPOT-E}

0 Citations
0 Influential
24 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!