2608.03875v1 Aug 04, 2026 cs.LG

구조 인지 미세 조정을 통한 VLM 보상 모델 성능 향상

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning

Xin Chen
Xin Chen
Citations: 93
h-index: 5
Pyrros Koussios
Pyrros Koussios
Citations: 0
h-index: 0
Andreas Krause
Andreas Krause
Citations: 63
h-index: 3
Chenhao Li
Chenhao Li
Citations: 127
h-index: 6

효과적인 보상 함수 설계는 강화 학습(RL)의 주요 난관 중 하나입니다. 최근 연구에서는 대규모 기초 시각-언어 모델(VLM)을 보상 모델로 활용하여, 텍스트와 관찰 결과 간의 유사성을 계산함으로써 수동적인 보상 설계 과정을 우회합니다. 하지만 이러한 보상은 종종 노이즈가 많고 신뢰성이 낮아, 실제 사용 시 직접적인 유용성이 제한되는 경우가 있습니다. 본 논문에서는 Structure-Aware Fine-Tuning (SAFT)이라는 간단하고 자기 지도 학습 기반의 방법을 제시하여, 외부 데이터 없이 온라인 환경에서 불완전한 보상 신호를 개선합니다. SAFT는 VLM의 잠재 공간을 LoRA 어댑터를 통해 정규화함으로써, 내재적인 구조적 선행 지식을 활용합니다. 우리는 다양한 수준의 기본 모델 능력을 가진 모델에 대한 엄격한 평가를 통해 SAFT의 다재다능함을 입증했습니다. 실험 결과, SAFT는 일관되게 보상 공간에서 노이즈를 줄여 정책 수렴 속도를 높이고, 기본 모델 대비 EPIC 거리 측면에서 상당한 성능 향상을 가져왔습니다. 이는 실패가 종종 의미론적 오해보다는 구조적인 취약성으로 인해 발생하는 경우가 많다는 것을 시사합니다. SAFT는 방대한 양의 인간 선호도 주석을 제거하고 작업에 내재된 구조적 유도 편향을 활용함으로써, 텍스트 기반 강화 학습을 안정화시키는 확장 가능한 방법을 제공하며, 작업 구조를 일반적인 유도 편향으로 통합하는 것의 더 넓은 가치를 강조합니다.

Original Abstract

Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward engineering. Although promising, these rewards are often noisy and unreliable, limiting their direct utility during deployment. We present Structure-Aware Fine-Tuning (SAFT), a simple, self-supervised method that refines these imperfect reward signals online without access to ground-truth supervision. SAFT leverages intrinsic structural priors to regularize the VLM's latent space via LoRA adapters. We rigorously evaluate SAFT across a spectrum of base model capabilities to demonstrate its versatility. Our results show that SAFT consistently denoises the reward landscape, yielding faster policy convergence and substantially improved alignment (EPIC distance) relative to the underlying base model, suggesting that failures can often be attributed to structural brittleness rather than semantic misunderstanding. By replacing extensive human preference annotation with structural inductive biases inherent to the task, SAFT offers a scalable path for stabilizing text-conditioned RL and underscores the broader value of incorporating task structure as a general inductive bias.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!