CineCap: 시공간적 기준점을 활용한 구조화된 추론을 통한 영화 영상 설명 생성
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
영화 영상 설명은 카메라 움직임, 샷 크기, 심도, 구도 및 촬영 각과 같은 전문적인 영화 용어를 사용하여 비디오가 어떻게 촬영되었는지 기술하는 것을 목표로 합니다. 이러한 기능은 세밀한 비디오 이해와 제어 가능한 고품질 비디오 생성에 중요하지만, 기존의 다중 모드 대규모 언어 모델에서는 아직 충분히 연구되지 않았습니다. 영화적 이해에 대한 질문-응답 기반 평가와 달리, 영화 영상 설명은 여러 영화적 차원에 걸쳐 일관된 개방형 설명을 요구합니다. 이 작업은 두 가지 주요 이유로 인해 어렵습니다. 첫째, 모델은 미묘한 시각적 증거로부터 전문적인 영화 용어를 추론해야 하며, 둘째, 모델은 포괄적이면서도 정확한 캡션을 생성해야 합니다. 이에 따라 우리는 구조화된 추론과 시공간적 기준점, 그리고 포괄성, 정확성 및 게이티드 커버리지 보상을 사용한 강화 학습을 결합한 프레임워크인 CineCap을 제안합니다. 전자는 전문적인 영화 설명을 명시적인 시각적 증거에 연결하고, 감독된 미세 조정(fine-tuning)을 위한 간결한 원자적 추론으로 구성합니다. 후자는 기술적 완전성과 사실 정확성 사이의 균형을 개선합니다. 또한 체계적인 평가를 위해 472개의 수동으로 주석이 달린 비디오-캡션 쌍으로 구성된 CineCap Bench라는 벤치마크를 구축했습니다. 광범위한 실험 결과, CineCap은 강력한 독점 및 오픈 소스 모델을 지속적으로 능가하며, 영화 영상 설명 분야에서 새로운 최고 성능을 달성하는 것으로 나타났습니다. 코드, 모델 체크포인트 및 벤치마크는 https://github.com/Hectormxy/CineCap.git 에서 공개적으로 이용할 수 있습니다.
Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and shooting angle. This capability is important for fine-grained video understanding and controllable movie-quality video generation, yet remains underexplored in existing multimodal large language models. Unlike question-answering-based evaluation of cinematic understanding, cinematographic captioning requires a unified open-form description over multiple cinematographic dimensions. This task is challenging for two main reasons: the model must infer professional cinematographic concepts from subtle visual evidence, and it must generate captions that are both comprehensive and accurate. Accordingly, we propose CineCap, a framework that combines structured reasoning with spatio-temporal anchors and reinforcement learning with comprehensiveness, accuracy, and gated coverage rewards. The former grounds professional cinematographic descriptions in explicit visual evidence and organizes them into compact atomic reasoning for supervised fine-tuning, while the latter improves the balance between descriptive completeness and factual correctness. In addition, we construct CineCap Bench, a benchmark of 472 manually annotated video-caption pairs for systematic evaluation. Extensive experiments show that CineCap consistently outperforms strong proprietary and open-source baselines, establishing a new state of the art for cinematographic captioning. The code, model checkpoint, and benchmark are publicly available in https://github.com/Hectormxy/CineCap.git.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.