AlayaWorld: 상호작용형 장기 예측 세계 모델링 - 기술 보고서 (v1.1)
AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
본 보고서는 AlayaWorld의 개선된 버전을 제시합니다. 이전 버전과 달리 핵심 아키텍처, 청크 단위 자기 회귀 생성 방식, 그리고 학습 데이터는 동일하게 유지되었지만, 조건 신호가 모델에 표현되고 통합되는 방식을 크게 수정했습니다. 새로운 설계는 간단한 원칙에 따라 이루어졌습니다: 조건 신호는 생성된 콘텐츠와 잠재적 표현 및 시간 구조 모두에서 최대한 밀접하게 일치해야 합니다. 이를 위해 두 가지 주요 변경 사항을 적용했습니다. 첫째, 기존의 깊이 변형 기반 공간 메모리를 스트리밍 3D 포인트 캐시 렌더러로 대체했습니다. 둘째, 시각적인 조건 신호가 동일한 인과 관계 VAE 잠재 공간에 인코딩되도록 조건 파이프라인을 재설계했으며, 생성된 비디오의 시간 통계와 일관성을 유지하도록 했습니다. 구체적으로, 새로운 버전은 다음과 같은 여섯 가지 수정 사항을 포함합니다: (1) 정적 프레임 이미지 조건을 모션 인식 잠재 조건으로 대체; (2) 재 렌더링된 공간 메모리를 연속적인 시퀀스로 인과적으로 인코딩; (3) 픽셀 공간에서의 시간 메모리 창을 정렬; (4) 메모리 토큰을 영으로 설정하는 대신 제거하는 하드 메모리 드롭아웃 채택; (5) 학습 및 추론 과정에서 VAE 인코딩 및 디코딩 프로토콜의 통합; 그리고 (6) 카메라 AdaLN 분기를 제거하여 시점 제어가 재 렌더링된 공간 조건만으로 제공되도록 변경.
This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure. To this end, we make two major changes. First, we replace the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer. Second, we redesign the conditioning pipeline so that visual conditions are encoded in the same causal-VAE latent space, with temporal statistics consistent with those of the generated video. Concretely, the new version introduces six modifications: (1) replacing static-frame image conditioning with motion-aware latent conditioning; (2) causally encoding re-rendered spatial memory as a continuous sequence; (3) aligning the temporal-memory window in pixel space; (4) adopting hard memory dropout that removes memory tokens rather than zeroing them; (5) unifying the VAE encoding and decoding protocol across training and inference; and (6) removing the camera AdaLN branch, such that viewpoint control is provided entirely through the re-rendered spatial condition.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.