2604.21718v1 Apr 22, 2026 cs.CV

인간-AI 감독을 통한 정밀 비디오 언어 구축

Building a Precise Video Language with Human-AI Oversight

Jiaxi Li
Jiaxi Li
Citations: 101
h-index: 6
Chancharik Mitra
Chancharik Mitra
Citations: 380
h-index: 6
Zhiqiu Lin
Zhiqiu Lin
Citations: 1,825
h-index: 15
Siyuan Cen
Siyuan Cen
Citations: 111
h-index: 3
Isaac Li
Isaac Li
Citations: 1
h-index: 1
Yuhang Huang
Yuhang Huang
Citations: 55
h-index: 2
Y. Ling
Y. Ling
Citations: 5
h-index: 1
Hewei Wang
Hewei Wang
Citations: 189
h-index: 6
Irene Pi
Irene Pi
Citations: 23
h-index: 2
Shihang Zhu
Shihang Zhu
Citations: 9
h-index: 1
R. Rao
R. Rao
Citations: 1
h-index: 1
George Liu
George Liu
Citations: 27
h-index: 2
Ruojin Li
Ruojin Li
Citations: 731
h-index: 12
Yi Han
Yi Han
Citations: 231
h-index: 5
Yilun Du
Yilun Du
Citations: 17
h-index: 2
D. Ramanan
D. Ramanan
Citations: 102,164
h-index: 95

비디오-언어 모델(VLMs)은 자연어를 통해 역동적인 시각 세계에 대한 추론을 학습합니다. 본 연구에서는 정밀한 비디오 캡셔닝을 가능하게 하는 확장 가능한 감독 시스템을 위한 공개 데이터셋, 벤치마크, 그리고 방법을 소개합니다. 먼저, 영화 제작자와 같은 전문가들이 개발한 수백 개의 정교하게 정의된 시각적 기본 요소에 기반하여, 피사체, 장면, 동작, 공간 관계, 카메라 움직임 및 카메라 동역학을 설명하는 체계적인 사양을 정의합니다. 다음으로, 고품질 캡션을 큐레이션하기 위해, 우리는 CHAI(Critique-based Human-AI Oversight)라는 프레임워크를 소개합니다. CHAI는 숙련된 전문가들이 모델이 생성한 초안 캡션을 비판하고 수정하여 개선된 최종 캡션을 생성하는 시스템입니다. 이러한 작업 분담은 모델이 텍스트 생성을 담당함으로써 주석 정확도와 효율성을 향상시키고, 인간이 검증에 더 집중할 수 있도록 합니다. 또한, 초안 및 최종 캡션 간의 비판 및 선호도는 SFT, DPO, 그리고 추론 시간 스케일링을 통해 오픈 소스 모델(Qwen3-VL)의 캡션 생성, 보상 모델링, 그리고 비판 생성 능력을 향상시키는 데 풍부한 감독 신호를 제공합니다. 우리의 실험 결과는, 우리의 감독 시스템에 의해 보장되는 정밀성, 재현율, 그리고 유용성 측면에서 비판 품질이 하위 작업의 성능에 직접적으로 영향을 미친다는 것을 보여줍니다. 비교적 적은 양의 전문가 감독을 통해, 결과 모델은 Gemini-3.1-Pro와 같은 독점 모델보다 뛰어난 성능을 보입니다. 마지막으로, 우리는 우리의 접근 방식을 사용하여 대규모 전문 비디오(예: 영화, 광고, 게임)를 다시 캡션하고, Wan과 같은 비디오 생성 모델을 세부적인 400단어 프롬프트를 더 잘 따르도록 미세 조정하여, 카메라 움직임, 각도, 렌즈, 초점, 시점, 그리고 프레이밍을 포함한 촬영술에 대한 더 정밀한 제어를 달성합니다. 우리의 결과는 정밀한 사양과 인간-AI 감독이 전문 수준의 비디오 이해 및 생성에 핵심적인 역할을 한다는 것을 보여줍니다. 데이터와 코드는 프로젝트 페이지에서 확인할 수 있습니다: https://linzhiqiu.github.io/papers/chai/

Original Abstract

Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable oversight that enable precise video captioning. First, we define a structured specification for describing subjects, scenes, motion, spatial, and camera dynamics, grounded by hundreds of carefully defined visual primitives developed with professional video creators such as filmmakers. Next, to curate high-quality captions, we introduce CHAI (Critique-based Human-AI Oversight), a framework where trained experts critique and revise model-generated pre-captions into improved post-captions. This division of labor improves annotation accuracy and efficiency by offloading text generation to models, allowing humans to better focus on verification. Additionally, these critiques and preferences between pre- and post-captions provide rich supervision for improving open-source models (Qwen3-VL) on caption generation, reward modeling, and critique generation through SFT, DPO, and inference-time scaling. Our ablations show that critique quality in precision, recall, and constructiveness, ensured by our oversight framework, directly governs downstream performance. With modest expert supervision, the resulting model outperforms closed-source models such as Gemini-3.1-Pro. Finally, we apply our approach to re-caption large-scale professional videos (e.g., films, commercials, games) and fine-tune video generation models such as Wan to better follow detailed prompts of up to 400 words, achieving finer control over cinematography including camera motion, angle, lens, focus, point of view, and framing. Our results show that precise specification and human-AI oversight are key to professional-level video understanding and generation. Data and code are available on our project page: https://linzhiqiu.github.io/papers/chai/

1 Citations
0 Influential
30 Altmetric
151.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!