2608.13546v1 Aug 13, 2026 cs.CV

Alaya-EVOKE: 선형 스케일링 감독 학습에서 무한 세계로

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Chuanhao Li
Chuanhao Li
Citations: 365
h-index: 9
Kaipeng Zhang
Kaipeng Zhang
Citations: 137
h-index: 5
Y. Zhan
Y. Zhan
Citations: 152
h-index: 6
Yuanyang Yin
Yuanyang Yin
Citations: 54
h-index: 4
Gongxuan Wang
Gongxuan Wang
Citations: 2
h-index: 1
Feng Zhao
Feng Zhao
Citations: 46
h-index: 4

상호작용 가능한 월드 모델은 지속적인 메모리, 즉각적인 반응성 및 장기적인 생성 능력을 지원해야 하지만, 이러한 요구 사항은 모델에 상충되는 제약을 가합니다. 디노이징 컨텍스트 또는 키-값 캐시에 과거 정보를 유지하는 것은 비용 증가를 초래하며, 세션 길이와 유지되는 메모리 간의 균형을 맞춰야 합니다. 또한, 낮은 지연 시간으로의 즉각적인 반응은 제한된 단계로 생성되므로, 그 능력은 '선생' 모델에 의해 제한됩니다. Evoke는 이러한 한계점을 극복하기 위해 지속적인 월드 상태를 외부화하고 장기적인 상호작용 생성을 위한 '선생' 모델을 재설계했습니다. 장면의 기하학적 정보는 카메라 인덱싱된 외부 월드 상태 저장소에 유지되며, 여기서 세션이 증가함에 따라 디노이징 컨텍스트가 제한되도록 관련 보기 정보만 검색됩니다. 우리는 '선생' 모델을 고정된 생성기로 취급하는 대신, 장기적인 감독 학습을 위해 설계했습니다. 희소 어텐션 메커니즘은 청크 단위 그룹화, 선택된 원격 프레임 검색 및 선형 어텐션 기반의 전역 상태를 결합하여 메모리와 연산량이 선형적으로 증가하면서도 장기적인 감독을 가능하게 합니다. 이러한 감독 학습은 짧은 시간 범위 내에서 지역적으로 타당성을 유지하는 콘텐츠 드리프트를 드러내며, 각 청크에 대한 조건부 설정은 시퀀스 전체에 걸쳐 프롬프트 변경 및 이벤트 제어를 가능하게 합니다. 30초 분량의 분포 일치 목표를 자가 강제 롤아웃 하에 적용하여 분류기-프리 가이드 없이 작동하는 세 단계 모델로 이러한 능력을 이전합니다. 이를 통해 장기적인 드리프트에 대한 저항성을 향상시키면서도 즉각적인 반응성 조건을 유지할 수 있습니다. 제한된 컨텍스트와 반복적인 외부 메모리를 통해 Evoke는 개방적이고 지속적으로 진화하는 생성을 지원합니다. 단일 H200 GPU에서 $384 imes 640$ 해상도의 각 1.5초 분량의 콘텐츠는 2.11초 만에 생성됩니다. 세 단계 월드 모델인 Evoke는 WBench에서 최첨단 성능을 달성했으며, VBench-Long 및 VBench-2.0에서도 경쟁력 있는 성능을 보입니다.

Original Abstract

Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at $384\times 640$, each $1.5\,\mathrm{s}$ chunk is generated in $2.11\,\mathrm{s}$. As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!