MuRA: 효율적이고 효과적인 테스트 시간 비전-언어 일반화를 위한 다중 순위 적응
MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization
비전-언어 모델은 놀라운 제로샷 능력을 보여주지만, 데이터 분포의 변화에 따라 성능 저하가 심각하게 나타납니다. 로우랭크 적응을 통한 테스트 시간 적응(TTA)은 파라미터 효율적인 해결책을 제공하지만, 현재 방법들의 근본적인 한계점은 정적인 순위 설정에 의존한다는 것입니다. 시각적 입력 데이터는 본질적으로 다양한 정보 밀도를 가지므로, 고정된 순위는 필연적인 최적화 타협으로 이어져 복잡한 장면에서는 과소 적합(underfitting)을 일으키고 단순한 장면에서는 과적합(overfitting)을 유발합니다. 이러한 격차를 해소하기 위해 우리는 토큰 수준의 시각적 복잡성에 따라 다양한 용량의 적응 모듈을 동적으로 선택하고 결합하는 새로운 프레임워크인 다중 순위 적응(MuRA)을 제안합니다. MuRA는 멀티랭크 직교 분해를 통해 우수한 지식 보존 초기화를 제공하며, 지속 가능한 의미-순위 매핑 학습을 위한 통합 구성 요소 융합과 연속 라우터 업데이트를 활용합니다. 또한, 우리는 이 적응 메커니즘의 필요성과 기울기 안정성을 수학적으로 증명하는 엄격한 이론적 근거를 제시합니다. 특히, MuRA의 동적인 설계는 가장 깊은 시각적 계층에서 작동하며, 가장 짧은 기울기 역전파 경로를 활용하여 뛰어난 성능을 발휘합니다. 광범위한 실험 결과는 MuRA가 다양한 도메인 일반화 및 교차 데이터셋 벤치마크에서 최첨단 정확도를 달성하는 동시에 계산 및 메모리 오버헤드를 크게 줄인다는 것을 보여줍니다.
Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank configurations. Because visual inputs inherently possess varying information densities, a fixed rank forces an inevitable optimization compromise, leading to underfitting on complex scenes and overfitting on simple ones. To bridge this gap, we propose Multi-Rank Adaptation (MuRA), a novel framework that dynamically selects and fuses adaptation modules of varying capacities based on token-level visual complexity. MuRA synergizes Multi-Rank Orthogonal Decomposition to provide a superior, knowledge-preserving initialization, and Unified Component Fusion with Continuous Router Updating to sustainably learn semantic-to-rank mappings. Furthermore, we provide rigorous theoretical justifications mathematically proving the necessity and gradient stability of this adaptive mechanism. Crucially, MuRA's dynamic design uniquely thrives at the deepest visual layer, capitalizing on the shortest gradient backpropagation path. Extensive experiments demonstrate that MuRA achieves state-of-the-art accuracy across extensive domain generalization and cross-dataset benchmarks while significantly reducing both computational and memory overhead.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.