2603.09160v1 Mar 10, 2026 cs.CV

RubiCap: 칭찬 기준 기반 강화 학습을 이용한 고밀도 이미지 캡셔닝

RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning

Tzu-Heng Huang
Tzu-Heng Huang
Citations: 99
h-index: 6
Sirajul Salekin
Sirajul Salekin
Citations: 103
h-index: 6
Javier Movellan
Javier Movellan
Citations: 5
h-index: 1
Frederic Sala
Frederic Sala
Citations: 57
h-index: 4
Manjot Bilkhu
Manjot Bilkhu
Citations: 60
h-index: 3

고밀도 이미지 캡셔닝은 시각-언어 사전 훈련 및 텍스트-이미지 생성에서 중요한 역할을 하지만, 전문가 수준의 어노테이션을 대량으로 확보하는 것은 매우 비쌉니다. 강력한 시각-언어 모델(VLMs)을 활용한 합성 캡셔닝은 실용적인 대안이지만, 지도 학습 기반의 지식 전달(distillation)은 종종 제한적인 출력 다양성과 낮은 일반화 성능을 보입니다. 강화 학습(RL)은 이러한 한계를 극복할 수 있지만, 지금까지는 검증 가능한 영역에서만 성공적인 결과를 보여주었으며, 이는 결정적인 검증 시스템에 의존합니다. 본 연구에서는 RubiCap이라는 새로운 RL 프레임워크를 통해 이러한 병목 현상을 해결합니다. RubiCap은 LLM(대규모 언어 모델)이 작성한 칭찬 기준(rubrics)으로부터 세밀하고 샘플별 보상 신호를 얻습니다. 먼저, RubiCap은 다양한 후보 캡션을 모으고, LLM 칭찬 기준 작성자를 활용하여 현재 정책의 강점과 부족한 점을 파악합니다. 이러한 정보는 명시적인 평가 기준으로 변환되어, LLM 평가자가 전체적인 품질을 세분화하여 평가하고, 단순한 스칼라 보상 대신 구조화되고 다면적인 평가를 수행하도록 합니다. 광범위한 벤치마크 테스트에서, RubiCap은 CapArena에서 가장 높은 성공률을 달성했으며, 지도 학습 기반 지식 전달, 기존 RL 방법, 전문가 어노테이션, 그리고 GPT-4V를 활용한 결과보다 우수한 성능을 보였습니다. CaptionQA에서는 더 높은 효율성을 보여주었습니다. 저희의 7B 모델은 Qwen2.5-VL-32B-Instruct 모델과 동등한 성능을 보이며, 3B 모델은 7B 모델보다 더 나은 성능을 보였습니다. 주목할 만한 점은, 작고 효율적인 RubiCap-3B 모델을 캡셔너로 사용하여 학습한 사전 훈련된 VLM이 독점 모델에서 생성된 캡션을 사용하여 학습한 모델보다 더 강력한 성능을 보인다는 것입니다.

Original Abstract

Dense image captioning is critical for cross-modal alignment in vision-language pretraining and text-to-image generation, but scaling expert-quality annotations is prohibitively expensive. While synthetic captioning via strong vision-language models (VLMs) is a practical alternative, supervised distillation often yields limited output diversity and weak generalization. Reinforcement learning (RL) could overcome these limitations, but its successes have so far been concentrated in verifiable domains that rely on deterministic checkers -- a luxury not available in open-ended captioning. We address this bottleneck with RubiCap, a novel RL framework that derives fine-grained, sample-specific reward signals from LLM-written rubrics. RubiCap first assembles a diverse committee of candidate captions, then employs an LLM rubric writer to extract consensus strengths and diagnose deficiencies in the current policy. These insights are converted into explicit evaluation criteria, enabling an LLM judge to decompose holistic quality assessment and replace coarse scalar rewards with structured, multi-faceted evaluations. Across extensive benchmarks, RubiCap achieves the highest win rates on CapArena, outperforming supervised distillation, prior RL methods, human-expert annotations, and GPT-4V-augmented outputs. On CaptionQA, it demonstrates superior word efficiency: our 7B model matches Qwen2.5-VL-32B-Instruct, and our 3B model surpasses its 7B counterpart. Remarkably, using the compact RubiCap-3B as a captioner produces stronger pretrained VLMs than those trained on captions from proprietary models.

5 Citations
2 Influential
3 Altmetric
24.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!