Wonder: 더욱 발전된 비디오 월드 모델
Wonder: Video World Model Done Better
본 논문에서는 실시간으로 카메라를 제어하며 환경을 탐험할 수 있는 범용 비디오 월드 모델인 Wonder를 소개합니다. Wonder는 주어진 이미지 또는 조건부 비디오로부터, 사용자가 인터랙티브하게 카메라를 움직여 새로운 영역을 발견하고, 이미 관찰했던 영역을 실시간으로 다시 방문하며 장기적인 시점에서 환경을 탐색할 수 있도록 작동하는 가상 세계를 구축합니다. 이러한 기능을 구현하기 위해서는 제어 방식, 메모리 메커니즘, 그리고 학습 전략의 시스템 수준에서의 통합 설계가 필요합니다. 우리는 모델이 카메라 움직임을 직접적으로 시각적 증거로 해석할 수 있도록 하는, 공간적으로 정렬된 운동 및 방향 정보를 제공하는 밀집 좌표 필드를 활용한 새로운 카메라 조건부 기법을 제시합니다. 또한, 증가하는 생성 맥락에서 빠르고 정확한 메모리 검색을 지원하기 위해 효율적인 희소 어텐션 기반 메모리 메커니즘을 제안합니다. 이를 통해 모델은 추론 시에 실제 맥락 길이에 관계없이 관련 있는 작은 세트의 컨텍스트 토큰에 선택적으로 집중할 수 있습니다. 더 나아가, 우리는 자기 강제 방식의 증류 파이프라인을 개선하기 위한 여러 기술을 개발하여, 학생 모델이 제어 신호를 잘 따르도록 하고 동시에 교사 모델로부터 다양한 생성 모드와 장기적인 메모리를 유지하도록 합니다. 이러한 구성 요소들을 결합하여 Wonder는 일관된 기하학적 구조, 외형, 그리고 역학적 특성을 유지하면서 16 FPS의 세밀한 비디오를 합성할 수 있습니다. 이미지-비디오 생성을 넘어, Wonder는 자연스럽게 비디오에 조건부로 제어되는 생성을 지원하여, 기존의 동적인 장면을 실시간으로 재촬영할 수 있도록 합니다.
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.