제로샷 불확실성을 존중하십시오: 테스트 시간 적응형 시각-언어 모델을 위한 보수적 교정
Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models
테스트 시간 적응(TTA)은 데이터 분포 변화 하에서 시각-언어 모델의 인식 정확도를 향상시킬 수 있지만, 종종 교정을 저하시켜 예측 신뢰성을 떨어뜨려 후속 의사 결정에 어려움을 초래합니다. 기존의 많은 라벨 없는 교정 방법은 프롬프트 최적화와 결합되거나, 예측 분포를 대략적으로만 설명하는 로짓 범위 통계에 의존합니다. 본 연구에서는 TTA가 최고 예측 정확도와 그 정확성이 변하지 않더라도 신뢰도를 높이고 엔트로피를 감소시킬 수 있다는 점을 보여줍니다. 이러한 현상을 '예측 보존형 샤프닝(prediction-preserving sharpening)'이라고 명명했습니다. 다양한 TTA 방법과 벤치마크에서, 쌍을 이루는 제로샷 예측에 비해 엔트로피 감소가 클수록 기대 교정 오차(ECE)가 증가하는 경향이 있습니다. 엔트로피가 감소된 샘플의 경우, 신뢰도 향상이 정확도 향상보다 큰 경향이 있습니다. 이러한 결과를 바탕으로, 본 연구에서는 각 샘플에 대한 불확실성 기준으로 제로샷 엔트로피를 사용하는 라벨 없는 후처리 방법인 '제로샷-앵커드 엔트로피 교정(ZAEC)'을 제안합니다. ZAEC은 최소한의 온도 스케일링을 통해 샤프닝된 예측의 제로샷 엔트로피를 선택적으로 복원하지만, 다른 모든 예측은 그대로 유지합니다. ZAEC은 라벨이 있는 교정 데이터나 학습 가능한 파라미터가 필요 없으며, 클래스 순위와 분류 정확도를 유지합니다. 5가지 TTA 방법과 15개 데이터 세트에 대해, ZAEC은 ViT-B/16 모델에서 가장 낮은 후처리 매크로 평균 ECE를 달성했으며, RN50 모델에서도 일관된 성능 향상을 보였습니다.
Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, making predictive confidence unreliable for downstream decision-making. Many existing label-free calibration approaches are either coupled to prompt optimization or rely on logit-range statistics that provide only a coarse characterization of the predictive distribution. We show that TTA can increase confidence and reduce entropy even when the top-1 prediction and its correctness remain unchanged, a failure mode we term prediction-preserving sharpening. Across diverse TTA methods and benchmarks, larger entropy reductions relative to paired zero-shot predictions are associated with greater increases in Expected Calibration Error (ECE). On entropy-reduced samples, confidence gains also tend to exceed accuracy gains. Based on these findings, we propose Zero-Shot-Anchored Entropy Calibration (ZAEC), a label-free post-hoc method that uses zero-shot entropy as a sample-specific uncertainty reference. ZAEC selectively restores the zero-shot entropy of sharpened predictions through minimal temperature scaling while leaving all other predictions unchanged. It requires no labeled calibration data or learned parameters and preserves class rankings and classification accuracy. Across five TTA methods and 15 datasets, ZAEC achieves the lowest post-hoc macro-average ECE on ViT-B/16, with consistent gains on RN50.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.