2607.17523v1 Jul 20, 2026 cs.CV

비디오로 생각하기: 비디오 생성 모델은 정말 현실 세계에 대해 추론할 수 있는가?

Thinking in Video: Can Video Generators Really Reason About the Real World?

Qiguang Chen
Qiguang Chen
SCIR
Citations: 1,830
h-index: 22
Libo Qin
Libo Qin
Citations: 708
h-index: 10
Yinghui Li
Yinghui Li
Citations: 89
h-index: 5
Di Yin
Di Yin
Citations: 3
h-index: 1
Xing Sun
Xing Sun
Citations: 94
h-index: 5
Yongheng Zhang
Yongheng Zhang
Citations: 258
h-index: 6
Yanchao Hao
Yanchao Hao
Citations: 30
h-index: 3
Zheng Wei
Zheng Wei
Citations: 6
h-index: 1
Ruihan Hou
Ruihan Hou
Citations: 0
h-index: 0
Xiaolong Liu
Xiaolong Liu
Citations: 0
h-index: 0
Guang Yang
Guang Yang
Citations: 0
h-index: 0
Ziang Liu
Ziang Liu
Citations: 108
h-index: 5
Manman Zhang
Manman Zhang
Citations: 0
h-index: 0
Hao Wu
Hao Wu
Citations: 26
h-index: 2
Peishan Dai
Peishan Dai
Citations: 0
h-index: 0

최근의 월드 모델 및 비디오 생성 기술 발전으로 인해, 비디오 생성 모델을 활용하여 실제 세계의 역학 관계를 시뮬레이션하고 예측하며 추론하는 새로운 패러다임이 등장했습니다. 우리는 이 패러다임을 '비디오로 생각하기(Thinking in Video)'라고 정의합니다. 여기서 비디오는 단순한 출력 결과물이 아니라, 인과적 사고를 구성하고 확장하며 검증하는 매개체입니다. 그러나 이러한 가능성은 아직 검증되지 않았습니다. 설득력 있는 결과물은 실제적인 이해가 아닌 단순히 암기된 특징을 반영할 수 있으며, 기존의 평가 지표는 시각적 충실도와 의미론적 논리를 분리합니다. 비디오 생성 모델이 이러한 추론 능력을 지원하는지 평가하기 위해, 우리는 두 가지 관점에서 월드 모델의 일관성을 감사하는 '인과-생성 이중 판별(Causal-Generative Dual-Judge, CGDJ)'를 도입했습니다. 명시적 인과 인식 테스트는 시공간적으로 압축된 비주얼 질의 응답을 통해 생성기가 비디오 시나리오를 추론 문제로 해석하는지 확인합니다. 반면, 암묵적 생성 인식-예측 격차(Implicit Generative Perception-Prediction Gap)는 생성기가 인과적인 결과를 일관된 미래 비디오로 표현하는지 평가합니다. 대표적인 오픈 소스 및 클로즈드 소스 생성 모델에 CGDJ를 적용한 결과, 명확한 인식-예측 격차가 나타났습니다. 오픈 소스 모델은 거의 0에 가까운 명시적 인과 인식에도 불구하고 그럴듯한 역학 관계를 생성하는 반면, 고급 클로즈드 소스 시스템은 더 강력하지만 여전히 제한적인 추론과 생성 간의 연관성을 보입니다. 추가 분석 결과, 오디오-비주얼 불일치가 드러났습니다. 모델은 비디오로 렌더링하는 것보다 음성으로 올바른 인과적 논리를 표현하는 데 더 안정적이며, 이는 '세계 시뮬레이터'라는 개념에 도전합니다.

Original Abstract

Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm as Thinking in Video, where video is not merely an output artifact but a medium for constructing, extending, and verifying causal thought. However, this promise remains unverified: convincing rollouts may reflect memorized appearances rather than causal understanding, while existing metrics separate perceptual fidelity from semantic logic. To evaluate whether video generators support such reasoning, we introduce the Causal-Generative Dual-Judge (CGDJ), auditing World Model Consistency from two perspectives. Explicit Causal Perception tests whether a generator reads a video scenario as a reasoning problem through spatio-temporal flattened visual question answering, while Implicit Generative Perception-Prediction Gap evaluates whether it renders the causal consequence as a consistent future video. Applying CGDJ to representative open- and closed-source generators reveals a clear Perception-Prediction Gap: open-source models produce plausible dynamics despite near-zero explicit causal perception, whereas advanced closed-source systems show stronger but still limited alignment between reasoning and generation. Further analysis exposes audio-visual misalignment, where models verbalize correct causal logic more reliably than they render it, challenging the "world simulator" narrative.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!