RefCaptioner: 다중 참조 이미지 기반 동영상 캡셔닝
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
기존의 동영상 캡셔닝 모델은 동영상 내용을 자연스럽게 설명하지만, 특정 시각적 요소를 여러 참조 이미지에 명시적으로 연결하지는 않습니다. 본 연구에서는 사실에 근거한 동영상 설명을 생성하고 문장 수준에서의 참조 이미지 연계를 요구하는 새로운 작업인 다중 참조 이미지 기반 동영상 캡셔닝을 소개합니다. 또한 이 작업을 위해 RefCaptioner라는 두 단계의 사후 학습 프레임워크를 제안합니다. RefCaptioner는 혼합 데이터 SFT(Supervised Fine-Tuning)와 계층적 커버리지-디스카운트 GRPO(Grounded Reinforcement Policy Optimization)를 결합하여 참조 이미지 선택, 문장 수준의 연계, 오답 제거 및 상호 참조 일관성을 동시에 향상시키면서 일반적인 동영상 캡셔닝 능력을 유지합니다. 학습을 지원하기 위해 $20,000$개의 동영상과 171,354개의 참조 이미지로 구성된 데이터셋을 구축했습니다. 또한 실제 동영상과 AI 생성 동영상을 모두 사용하여 캡션의 사실성 및 다중 참조 연계를 평가하는 MRVBench라는 새로운 벤치마크를 소개합니다. 실험 결과, RefCaptioner는 공개 모델 중에서 가장 우수한 전반적인 성능을 보였으며, 표준 동영상 캡셔닝 벤치마크에서도 경쟁력 있는 성능을 유지했습니다. 사용자 평가 결과, RefCaptioner가 생성한 캡션이 평가자들에게 더 선호되었으며, 오픈 소스 및 독점 동영상 생성기로 더 정확하게 동영상을 재구성하는 데 도움이 된다는 사실이 확인되었습니다.
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing $20,000$ videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.