테스트 시점에서의 대규모 시각-언어 모델 정렬: 경로 기반 구조화된 샘플링 접근 방식
Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach
대규모 시각-언어 모델(LVLM)을 인간의 의도 및 시각적 추론 작업의 요구 사항에 맞추기 위해 일반적으로 훈련 후 강화 학습(RL) 알고리즘이 사용됩니다. 그러나 기존의 RL 기반 정렬 방법은 종종 많은 자원을 필요로 하며, 훈련 목표와 추론 시간 분포 간의 불일치를 경험하는 경우가 많습니다. 이러한 격차를 해소하기 위해, 우리는 동적 추론 시간 개선을 위한 경로 기반 구조화된 샘플링을 활용하는 새로운 테스트 시점 정렬 접근 방식을 제안합니다. 이를 통해 시각적 연결 및 논리적 일관성을 향상시킵니다. 우리의 접근 방식은 먼저 경로 학습 알고리즘을 사용하여 추론 메모리 뱅크를 구축하며, 이는 복잡한 질문 해결을 미리 정의된 추론 패턴의 정렬된 순서로 분해합니다. 그 후, 추론 메모리 뱅크에서 경로를 수집하여 글로벌 구조적 추론 사전 지식을 설정하고, 이어서 반복적인 마르코프 체인 몬테카를로(MCMC) 알고리즘을 사용하여 추론 과정을 지역적으로 다중 목표 최적화합니다. 여러 멀티모달 추론 데이터 세트에 대한 실험 결과는 우리의 접근 방식이 상당한 정확도 향상을 가져오면서도 과도한 추론 오버헤드를 발생시키지 않음을 보여줍니다. 이러한 결과는 경로 기반 테스트 시점 샘플링을 전통적인 훈련 후 정렬의 확장 가능하고 효과적인 대안으로 확립하며, 특히 복잡한 시각적 추론 작업에 유용합니다.
Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks. However, existing RL-based alignment methods are often resource-intensive and encounter mismatches between training objectives and inference-time distributions. To bridge this gap, we propose a novel test-time alignment approach that leverages trajectory-guided structured sampling for dynamic inference-time refinement, achieving better alignment with visual grounding and ensuring logical consistency. Our approach begins with curating a reasoning memory bank via a trajectory learning algorithm, which decomposes complex question solving into ordered sequences of predefined reasoning patterns. It subsequently accomplishes inference-time alignment by first collecting trajectories from reasoning memory bank to establish a global structural reasoning prior, and then using an iterative Markov Chain Monte Carlo (MCMC) algorithm for localized multi-objective refinement of the reasoning trace. Experiments across multiple multimodal reasoning datasets demonstrate that our approach significantly improves accuracy without incurring prohibitive inference overhead. These results establish trajectory-guided test-time sampling as a scalable and effective alternative to traditional post-training alignment, particularly for complex visual reasoning tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.