2605.28023v1 May 27, 2026 cs.CV

VCap: 약한 모델에서 강한 모델로의 시각 설명 생성: 초기분포 기반 보상

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

Yankai Yang
Yankai Yang
Citations: 22
h-index: 3
Yancheng Long
Yancheng Long
Citations: 13
h-index: 2
Tianke Zhang
Tianke Zhang
Citations: 351
h-index: 8
Kaiyu Jiang
Kaiyu Jiang
Citations: 281
h-index: 5
Haonan Fan
Haonan Fan
Citations: 175
h-index: 3
Changyi Liu
Changyi Liu
Citations: 317
h-index: 7
Tingting Gao
Tingting Gao
Citations: 503
h-index: 12
Xingyu Lu
Xingyu Lu
Citations: 140
h-index: 3
Jinpeng Wang
Jinpeng Wang
Citations: 595
h-index: 12
Xuanyu Zheng
Xuanyu Zheng
Citations: 13
h-index: 1
Bin Wen
Bin Wen
Citations: 489
h-index: 10
Hanqi Li
Hanqi Li
Citations: 44
h-index: 4
Yiyang Fan
Yiyang Fan
Citations: 346
h-index: 2
Chun Yuan
Chun Yuan
Citations: 88
h-index: 3
Yi-Fan Zhang
Yi-Fan Zhang
Citations: 137
h-index: 3
Fan Yang
Fan Yang
Citations: 140
h-index: 6

시각 설명 생성은 모델이 시각적 내용을 정확하게 반영하면서 누락과 환각을 최소화해야 합니다. 현재 주류 패러다임인 대규모 언어 모델(MLLM)은 확장 및 고품질 데이터를 통해 강력한 성능을 달성했습니다. 최근 강화 학습(RL)은 MLLM의 정밀도와 범위를 향상시키는 중요한 방법으로 부상했지만, 기존의 시각 설명 생성 보상 설계는 사실 확인에 필요한 세분화되고 신뢰할 수 있는 정보를 제공하지 못하여 효과가 제한적입니다. 이러한 문제를 해결하기 위해 우리는 VCap을 제안합니다. VCap은 참조 캡션(증거)과 시각 정보(심판)를 결합하는 증인-심판 보상 시스템입니다. VCap은 시각 정보에 기반한 정책이 생성한 캡션과 참조 캡션 간의 사실적 일관성을 명시적으로 검증함으로써, 캡션 품질 검증을 위한 초기분포 수준의 정밀도를 가진 보상 신호를 제공합니다. 이러한 설계는 불완전한 참조 데이터에서도 효과적인 학습을 가능하게 하며, RL 학습에서 약한 모델에서 강한 모델로의 일반화 능력을 향상시킵니다. 실험 결과, VCap으로 훈련된 80억 파라미터 모델은 다양한 이미지 및 비디오 설명 벤치마크에서 공개 및 비공개 최고 성능 모델보다 우수한 성능을 보였습니다. 인간 평가를 통해 사실 정확도와의 높은 일관성이 다시 한번 확인되었습니다. 또한, VCap은 MLLM의 인지 능력을 향상시키고, 다양한 작업에 대한 일반화 능력을 높이며, N-Best 증류 기법보다 뛰어난 성능을 보여주어 기존의 RLVR에 대한 선입견에 도전합니다.

Original Abstract

Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for captioning, MLLMs have achieved strong performance through scaling and high-quality data. Recently, RL has emerged as a key route to driving MLLMs toward higher precision and broader coverage, however, existing reward designs for captioning fail to provide fine-grained and reliable signals for factual verification, limiting their effectiveness. To address this, we propose VCap, a Witness-Adjudicator reward that pairs the reference caption (a witness) with the visual signal (an adjudicator). By explicitly verifying factual consistency between the reference and policy-generated captions grounded in the visual signal, VCap delivers a reward signal with hypergeometric-distribution-level precision for caption quality verification. This design enables effective learning even from imperfect references, facilitating weak-to-strong generalization in RL training. In our experiments, an 8B model trained with VCap outperforms open- and closed-source SOTA models on multiple image and video captioning benchmarks. Human evaluation further confirms its strong alignment with factual correctness. Additionally, VCap improves MLLM perceptual capability, generalizes across tasks, and surpasses best-of-N distillation, challenging prior assumptions about RLVR.

1 Citations
0 Influential
6 Altmetric
31.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!