2602.21655v1 Feb 25, 2026 cs.CV

CCCaption: 완전하고 정확한 이미지 캡셔닝을 위한 이중 보상 강화 학습

CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning

Zhijiang Tang
Zhijiang Tang
Citations: 12
h-index: 2
Linhua Wang
Linhua Wang
Citations: 9
h-index: 1
Peng Hou
Peng Hou
Citations: 3
h-index: 1
Anxiang Zeng
Anxiang Zeng
Citations: 11
h-index: 1
Jianqiang Huang
Jianqiang Huang
Citations: 7
h-index: 2
Jiaxin Qi
Jiaxin Qi
Citations: 15
h-index: 2
Weihao Jiang
Weihao Jiang
Citations: 14
h-index: 2

이미지 캡셔닝은 시각-언어 이해의 기본적인 과제이지만, 여전히 정확한 지침은 주로 사람이 직접 작성한 참고 자료에 의존합니다. 사람의 주석은 주관적인 선호와 전문성을 반영하기 때문에, 정확한 캡션은 종종 불완전하거나 심지어 부정확하며, 이는 캡션 모델의 성능을 제한합니다. 우리는 캡션의 품질이 완전성(캡션이 이미지의 중요한 시각적 사실을 모두 포함하는가?)과 정확성(설명이 이미지와 일치하는가?)이라는 두 가지 객관적인 측면으로 평가되어야 한다고 주장합니다. 이러한 점을 염두에 두고, 우리는 CCCaption을 제안합니다. CCCaption은 특정 이미지 캡셔닝 데이터셋을 사용하여 훈련된, 이중 보상 강화 학습 프레임워크로, 이미지의 extbf{C}omplete (완전)하고 extbf{C}orrect (정확)한 extbf{Captions} (캡션)을 생성하기 위해 이러한 특성을 명시적으로 최적화합니다. 완전성을 위해, 우리는 다양한 언어 모델(LVLM)을 사용하여 이미지를 여러 시각적 질의로 분해하고, 더 많은 질의에 답하는 캡션에 보상을 제공하며, 훈련 효율성을 향상시키기 위해 동적 질의 샘플링 전략을 사용합니다. 정확성을 위해, 우리는 캡션 분해에서 파생된 부분 캡션 질의의 진실성을 검증하여 환각을 포함하는 캡션에 페널티를 부여합니다. 우리의 대칭적인 이중 보상 최적화는 완전성과 정확성을 동시에 극대화하여 모델이 이러한 객관적인 기준을 더 잘 충족하는 캡션을 생성하도록 유도합니다. 표준 캡셔닝 벤치마크에 대한 광범위한 실험 결과, 일관된 성능 향상이 확인되었으며, 이는 사람이 작성한 참고 자료를 단순히 모방하는 것을 넘어 캡션 모델을 훈련하는 원칙적인 방법을 제시합니다.

Original Abstract

Image captioning remains a fundamental task for vision language understanding, yet ground-truth supervision still relies predominantly on human-annotated references. Because human annotations reflect subjective preferences and expertise, ground-truth captions are often incomplete or even incorrect, which in turn limits caption models. We argue that caption quality should be assessed by two objective aspects: completeness (does the caption cover all salient visual facts?) and correctness (are the descriptions true with respect to the image?). To this end, we introduce CCCaption: a dual-reward reinforcement learning framework with a dedicated fine-tuning corpus that explicitly optimizes these properties to generate \textbf{C}omplete and \textbf{C}orrect \textbf{Captions}. For completeness, we use diverse LVLMs to disentangle the image into a set of visual queries, and reward captions that answer more of these queries, with a dynamic query sampling strategy to improve training efficiency. For correctness, we penalize captions that contain hallucinations by validating the authenticity of sub-caption queries, which are derived from the caption decomposition. Our symmetric dual-reward optimization jointly maximizes completeness and correctness, guiding models toward captions that better satisfy these objective criteria. Extensive experiments across standard captioning benchmarks show consistent improvements, offering a principled path to training caption models beyond human-annotation imitation.

3 Citations
0 Influential
1 Altmetric
8.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!