2608.04964v1 Aug 05, 2026 cs.AI

WorldCycle: 장기 비디오 세계 모델을 위한 자체 검증 강화 학습

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

Bohai Gu
Bohai Gu
Citations: 37
h-index: 2
Taiyi Wu
Taiyi Wu
Citations: 4
h-index: 1
Dazhao Du
Dazhao Du
Citations: 378
h-index: 8
Jian Liu
Jian Liu
Citations: 345
h-index: 3
Xiaotong Zhao
Xiaotong Zhao
Citations: 133
h-index: 4
Alan Zhao
Alan Zhao
Citations: 29
h-index: 3
Jie Zhang
Jie Zhang
Citations: 60
h-index: 5
Xiaoyi Pang
Xiaoyi Pang
Citations: 6
h-index: 2
Song Guo
Song Guo
Citations: 91
h-index: 3
Xiaocheng Lu
Xiaocheng Lu
Citations: 358
h-index: 9
Haobin Zhong
Haobin Zhong
Citations: 45
h-index: 1
Yueyang Yuan
Yueyang Yuan
Citations: 12
h-index: 2

인터랙티브 비디오 세계 모델은 장기 계획 및 탐색에 필수적이지만, 누적되는 오류로 인해 어려움을 겪습니다. 강화 학습(RL)과 같은 사후 훈련 방법은 이러한 모델의 성능을 향상시킬 수 있지만, 검증이라는 난관에 부딪힙니다. 임의의 행동 시퀀스에 대해 장기적인 추세를 측정할 수 있는 ground truth 미래 상태가 존재하지 않기 때문입니다. 저희 연구의 핵심 아이디어는 가역적인 행동 주기가 이러한 검증을 가능하게 한다는 것입니다. 역순으로 구성된 행동 시퀀스는 분석적으로 초기 상태로 돌아와야 하며, 이를 통해 레이블이 없는 방식으로 장기적인 정확성을 감독할 수 있습니다. 이러한 점을 바탕으로, 저희는 자체 검증 강화 학습 프레임워크인 WorldCycle을 소개합니다. WorldCycle은 일반적인 행동 시퀀스에서 폐쇄된 행동 주기를 구성하고 반복 실행하며, 두 가지 상호 보완적인 보상을 최적화합니다. 첫 번째는 거울상으로 반전된 순방향 및 역방향 세그먼트 간의 대칭성을 강제하는 공간적 폐쇄 보상이고, 두 번째는 반복되는 주기 실행에서 상태를 정렬하는 시간적 일관성 보상입니다. 이러한 보상은 모델이 단순히 기억한 시간 패턴 대신 일관된 상태 연산자로 행동을 학습하도록 유도하며, 기본 모델이 제대로 처리하지 못하는 분포 외의 복합적인 행동 주기로 자연스럽게 확장됩니다. 또한, 상태 반환 능력을 진단하기 위한 벤치마크인 CycleBench을 공개합니다. WorldCycle은 상태 반환 드리프트를 최대 44%까지 줄이고, 복합적인 행동 정확도를 기본 모델보다 거의 4배 향상시켜 물리적으로 기반한 세계 모델의 중요한 토대를 제공합니다.

Original Abstract

Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary action sequences, no ground-truth future state exists to measure long-term drift. Our key insight is that reversible action cycles make this verification possible: a sequence composed with its inverse must analytically return to the initial state, yielding annotation-free supervision on long-horizon correctness. Building on this, we introduce WorldCycle, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions. These rewards force the model to learn actions as consistent state operators rather than memorized temporal patterns, and extend naturally to out-of-distribution composite action cycles that the base model handles poorly. We further release CycleBench, a diagnostic benchmark for state-returning ability under complex action structures. WorldCycle reduces state returning drift by up to 44% and lifts composite-action accuracy nearly 4x over the base model, providing a vital foundation for physically grounded world models.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!