2607.08763v1 Jul 09, 2026 cs.CV

OpenCoF: 비디오 생성을 통한 추론 학습

OpenCoF: Learning to Reason Through Video Generation

Ziyu Guo
Ziyu Guo
Citations: 2,663
h-index: 20
Renrui Zhang
Renrui Zhang
Citations: 418
h-index: 9
Dongzhi Jiang
Dongzhi Jiang
Citations: 1,565
h-index: 14
Xinyan Chen
Xinyan Chen
Citations: 202
h-index: 4
Hongsheng Li
Hongsheng Li
Citations: 1,825
h-index: 16

추론은 특히 논리적인 결과를 이해하여 신뢰할 수 있는 결정을 내려야 하는 대규모 모델에게 있어 핵심적인 역량이 되었습니다. 최근의 비디오 생성 모델들은 이전의 체인 오브 씽크(Chain-of-Thought, CoT) 방식과는 다른 추론 경로를 제공합니다. 이 경로는 시간적으로 연결된 프레임을 통해 전개되며, 이를 체인 오브 프레임(Chain-of-Frame, CoF) 추론이라고 합니다. 그러나 기존의 비디오 생성 모델들은 주로 일반적인 비디오 데이터셋으로 훈련되어 있으며, 여전히 다양한 지도 학습과 CoF 추론을 위한 특화된 설계가 부족합니다. 이러한 격차를 해결하기 위해, 우리는 OpenCoF-17K 데이터셋(11개의 작업 범주를 포괄하는 추론 비디오 데이터셋)과 Wan-CoF(다양한 시간적 감독이 CoF 행동을 향상시키는지 연구하기 위한 미세 조정된 비디오 모델)로 구성된 프레임워크인 OpenCoF를 소개합니다. 네 가지의 비디오 추론 벤치마크에서, Wan-CoF는 Wan2.2-I2V-A14B 기준 모델보다 상당한 성능 향상을 보였습니다. 이를 바탕으로, 우리는 CoF 능력을 위한 더욱 발전된 설계를 경험적으로 탐구했습니다. 즉, 시각적 및 텍스트 추론 토큰을 사용하여 모델을 구성했습니다. 이 메커니즘은 각각 저수준의 시각적 단서와 고수준의 의미론적 사전 지식을 공간적 및 시간적 추론에 활용합니다. 성능 비교 및 어텐션 분석을 통해, 이러한 토큰들이 모델 깊이, 디노이징 단계, 공간 및 시간에 걸쳐 어떻게 기여하는지 살펴보았습니다. 우리의 결과는 강력한 비디오 추론이 광범위한 시간적 감독과 중간 추론 상태를 구성하기 위한 명시적인 메커니즘 모두를 필요로 한다는 것을 시사합니다. 우리는 데이터셋, 모델 및 코드를 공개하여 추론 지향적인 비디오 생성에 대한 미래 연구를 촉진하고자 합니다.

Original Abstract

Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold through temporally connected frames, known as Chain-of-Frame (CoF) reasoning. However, existing video generators are primarily trained on general video corpora, still lacking diverse supervision and dedicated designs for CoF reasoning. To address this gap, we introduce OpenCoF, a framework comprising the OpenCoF-17K dataset, a reasoning video dataset spanning 11 task families, and Wan-CoF, a fine-tuned video model for studying whether diverse temporal supervision improves CoF behavior. Across four video reasoning benchmarks, Wan-CoF achieves considerable gains over the Wan2.2-I2V-A14B baseline. Building on this, we empirically explore more advanced designs for CoF capabilities, i.e., equipping the model with visual and textual reasoning tokens. This mechanism respectively captures low-level visual cues and high-level semantic priors for spatial and temporal reasoning. Through performance comparisons and attention analysis, we examine how these tokens contribute across model depth, denoising steps, space, and time. Our results suggest that stronger video reasoning requires both broad temporal supervision and explicit mechanisms for organizing intermediate reasoning state. We open-source the dataset, model, and code to facilitate future research on reasoning-oriented video generation.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!