2608.02216v1 Aug 03, 2026 cs.CV

비전-언어 모델의 테스트 시간 적응을 위한 로컬 마진 복원

Local Margin Restoration for Test-Time Adaptation of Vision-Language Models

Guowei Wang
Guowei Wang
Citations: 14
h-index: 2
Xin Lin
Xin Lin
Citations: 664
h-index: 9
Yan Huang
Yan Huang
Citations: 2
h-index: 1
Kangjun Liu
Kangjun Liu
Citations: 75
h-index: 4
Xu Wang
Xu Wang
Citations: 0
h-index: 0

CLIP과 같은 비전-언어 모델(VLM)은 뛰어난 제로샷 성능을 보여주지만, 예상치 못한 테스트 데이터 분포 변화가 발생하면 성능이 급격하게 저하되는 경우가 많습니다. 테스트 시간 적응(TTA)은 이러한 문제를 해결할 수 있는 유망한 방법이지만, 레이블이 없는 테스트 데이터 스트림에 대해 VLM을 지속적으로 적응시키는 것은 근본적인 어려움을 야기합니다. 기존의 상위 1개 결과 중심 업데이트는 관련 클래스 간의 로컬 의미론적 구조를 손상시켜 오류를 강화하는 경향이 있으며, 반복적인 적응은 점진적인 편향 누적을 심화시켜 결국 모델을 모드 콜랩스로 몰아넣습니다. 이러한 문제점을 해결하기 위해, 우리는 가볍고 단계가 하나인 TTA 프레임워크인 로컬 마진 복원(LMR)을 제안합니다. LMR은 샘플 수준에서 보호된 마진 복원(PMR)이라는 목적 함수를 사용하여 외부의 어려운 부정 예시로부터 가능성이 높은 상위 후보들을 보호함으로써 로컬 의미론적 구조를 복구합니다. 동시에, 스트림 수준에서의 성능 저하를 방지하기 위해, 적응형 마진(AM) 컨트롤러와 편향 보정(BC)이라는 이중 단계 안정화 메커니즘을 도입하여 점진적인 편향 누적을 동적으로 파괴하고 모드 콜랩스를 예방합니다. CIFAR-C, ImageNet-C 및 ImageNet 변형 데이터셋에 대한 광범위한 실험 결과, LMR은 최첨단 TTA 방법보다 일관되게 우수한 성능을 보이며, 특히 배치 크기가 작은 어려운 테스트 환경에서도 뛰어난 강건성과 효율성을 입증했습니다. 저희의 코드는 다음 링크에서 확인할 수 있습니다: https://github.com/DennisHuangYan/LMR.

Original Abstract

Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected test-time distribution shifts. While Test-Time Adaptation (TTA) offers a promising solution, continuously adapting VLMs over an unlabeled test stream presents fundamental challenges. Conventional top-1-centric updates often reinforce errors by corrupting the local semantic geometry among related classes, while iterative adaptation exacerbates progressive bias accumulation, ultimately driving the model toward mode collapse. To overcome these coupled vulnerabilities, we propose Local Margin Restoration (LMR), a lightweight, one-step TTA framework. At the sample level, our Protected Margin Restoration (PMR) objective recovers local semantic geometry by shielding plausible near-top candidates from external hard negatives. Concurrently, to combat stream-level degradation, we introduce a dual-stage stabilization mechanism, featuring an Adaptive Margin (AM) controller and Bias Correction (BC), to dynamically disrupt progressive bias accumulation and prevent mode collapse. Extensive experiments on CIFAR-C, ImageNet-C, and ImageNet variants demonstrate that LMR consistently outperforms state-of-the-art TTA baselines, proving exceptionally robust and efficient even in challenging low-batch test-time regimes. Our code is available at https://github.com/DennisHuangYan/LMR.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!