AVCap: 세부 정보 인지 보상을 활용한 오디오-비디오 통합 자막 강화
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
세밀한 오디오-비디오 통합 자막은 다중 모드 비디오 이해 및 생성에 필수적입니다. 그러나 기존 연구는 다음과 같은 세 가지 주요 한계점을 가지고 있습니다: (1) 미세하게 조정된 오디오-비디오 통합 자막을 포함하는 고품질 공개 데이터셋의 부족, (2) 조잡한 보상 신호에 의존하는 강화 학습 방법, 그리고 (3) 원자 수준에서 상세한 오디오-비디오 자막을 평가하기 위한 벤치마크 및 지표의 부재. 이러한 문제점을 해결하기 위해 우리는 다음과 같은 내용을 제안합니다: (1) AVCap-100K, 시간적으로 정렬되고 세부적인 정보를 담은 100만 건의 오디오-비디오 자막으로 구성된 고품질 데이터셋, (2) Detail-Aware GRPO (Da-GRPO)를 통해 최적화되어 오픈 소스 모델 중 최고 성능을 달성하고 여러 평가에서 독점 모델과 동등하거나 능가하는 AVCap 모델, 그리고 (3) 오디오-비디오 자막의 원자 수준 세부 정보를 평가하기 위한 특수 벤치마크 및 지표인 AVCap-Bench와 AVCap-Score. 저희 코드, 모델, 데이터셋은 https://huggingface.co/collections/Apryle/avcap 에서 이용하실 수 있습니다.
Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main limitations: (1) the scarcity of high-quality public datasets with fine-grained audio-visual joint captions; (2) reinforcement-learning methods that rely on coarse reward signals; and (3) the lack of a benchmark and metric for evaluating detailed audiovisual captions at the atomic level. To address these challenges, we propose: (1) AVCap-100K, a high-quality dataset of 100K temporally aligned, detail-rich audio-video captions; (2) AVCap, a model optimized via Detail-Aware GRPO (Da-GRPO) that achieves state-of-the-art performance among open-source models and matches or surpasses proprietary models on several evaluations; and (3) AVCap-Bench and AVCap-Score, a specialized benchmark and metric for evaluating atomic-level details in audiovisual captions. Our code, models, and datasets are available at https://huggingface.co/collections/Apryle/avcap.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.