2608.13492v1 Aug 13, 2026 cs.AI

AlayaWorld: 상호작용형 장기 예측 세계 모델링 - 기술 보고서 (v1.1)

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

Zihui Gao
Zihui Gao
Citations: 32
h-index: 2
Chuanhao Li
Chuanhao Li
Citations: 365
h-index: 9
Kang He
Kang He
Citations: 14
h-index: 2
Zian Meng
Zian Meng
Citations: 2
h-index: 1
Y. Zhan
Y. Zhan
Citations: 152
h-index: 6
Liaoyuan Fan
Liaoyuan Fan
Citations: 25
h-index: 4
Xuangeng Chu
Xuangeng Chu
Citations: 31
h-index: 3
Ruicong Liu
Ruicong Liu
Citations: 57
h-index: 4
Yuanyang Yin
Yuanyang Yin
Citations: 54
h-index: 4
AlayaWorld Team Kaipeng Zhang
AlayaWorld Team Kaipeng Zhang
Citations: 0
h-index: 0
Yongtao Ge
Yongtao Ge
Citations: 429
h-index: 7
Jiaming Tan
Jiaming Tan
Citations: 1
h-index: 1
Mingliang Zhai
Mingliang Zhai
Citations: 34
h-index: 3
Xiaojie Xu
Xiaojie Xu
Citations: 35
h-index: 2
Zhen Li
Zhen Li
Citations: 165
h-index: 5
Zhe Lin
Zhe Lin
Citations: 14
h-index: 2
Zhixiang Wang
Zhixiang Wang
Citations: 5
h-index: 2

본 보고서는 AlayaWorld의 개선된 버전을 제시합니다. 이전 버전과 달리 핵심 아키텍처, 청크 단위 자기 회귀 생성 방식, 그리고 학습 데이터는 동일하게 유지되었지만, 조건 신호가 모델에 표현되고 통합되는 방식을 크게 수정했습니다. 새로운 설계는 간단한 원칙에 따라 이루어졌습니다: 조건 신호는 생성된 콘텐츠와 잠재적 표현 및 시간 구조 모두에서 최대한 밀접하게 일치해야 합니다. 이를 위해 두 가지 주요 변경 사항을 적용했습니다. 첫째, 기존의 깊이 변형 기반 공간 메모리를 스트리밍 3D 포인트 캐시 렌더러로 대체했습니다. 둘째, 시각적인 조건 신호가 동일한 인과 관계 VAE 잠재 공간에 인코딩되도록 조건 파이프라인을 재설계했으며, 생성된 비디오의 시간 통계와 일관성을 유지하도록 했습니다. 구체적으로, 새로운 버전은 다음과 같은 여섯 가지 수정 사항을 포함합니다: (1) 정적 프레임 이미지 조건을 모션 인식 잠재 조건으로 대체; (2) 재 렌더링된 공간 메모리를 연속적인 시퀀스로 인과적으로 인코딩; (3) 픽셀 공간에서의 시간 메모리 창을 정렬; (4) 메모리 토큰을 영으로 설정하는 대신 제거하는 하드 메모리 드롭아웃 채택; (5) 학습 및 추론 과정에서 VAE 인코딩 및 디코딩 프로토콜의 통합; 그리고 (6) 카메라 AdaLN 분기를 제거하여 시점 제어가 재 렌더링된 공간 조건만으로 제공되도록 변경.

Original Abstract

This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure. To this end, we make two major changes. First, we replace the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer. Second, we redesign the conditioning pipeline so that visual conditions are encoded in the same causal-VAE latent space, with temporal statistics consistent with those of the generated video. Concretely, the new version introduces six modifications: (1) replacing static-frame image conditioning with motion-aware latent conditioning; (2) causally encoding re-rendered spatial memory as a continuous sequence; (3) aligning the temporal-memory window in pixel space; (4) adopting hard memory dropout that removes memory tokens rather than zeroing them; (5) unifying the VAE encoding and decoding protocol across training and inference; and (6) removing the camera AdaLN branch, such that viewpoint control is provided entirely through the re-rendered spatial condition.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!