2608.07408v1 Aug 07, 2026 cs.CV

비디오 월드 모델을 위한 어드레서블 메모리

Addressable Memory for Video World Models

Xindi Wu
Xindi Wu
Citations: 185
h-index: 7
Despoina Paschalidou
Despoina Paschalidou
Citations: 677
h-index: 2
Laura Leal-Taix'e
Laura Leal-Taix'e
Citations: 11
h-index: 2
Olga Russakovsky
Olga Russakovsky
Citations: 1,454
h-index: 15
Jonathan Lorraine
Jonathan Lorraine
NVIDIA
Citations: 1,454
h-index: 13
Sven Elflein
Sven Elflein
Citations: 79
h-index: 4
James Lucas
James Lucas
Citations: 138
h-index: 5
Aljosa Osep
Aljosa Osep
Citations: 4,407
h-index: 24

본 논문에서는 대화형 비디오 월드 모델에서의 시각적 지속성을 연구합니다. 이러한 모델은 이전에 생성된 프레임을 저장하기 위해 키-값(KV) 캐시를 활용하는 방식으로 작동하며, 이는 일종의 시각적 메모리 역할을 합니다. 그러나 훈련 데이터 범위를 벗어난 긴 시퀀스에서 모델이 저장된 콘텐츠에 정확하게 접근하는 데 어려움을 겪는다는 점을 발견했습니다. 그 이유는 시간 기반 로터리 포지셔널 임베딩(RoPE) 오프셋이 훈련 중에 관찰되었던 범위를 벗어나기 때문이며, 이로 인해 모델은 어텐션을 통해 관련 시각 정보를 정확하게 검색하는 데 어려움을 겪습니다. 또한, RoPE-회전 공간에서 캐시를 무분별하게 압축하면 위치 정보가 손상되어 호환되지 않는 위치 단계들이 평균화되기 때문입니다. 이러한 문제를 해결하기 위해, 우리는 장기적인 시각적 지속성을 위한 훈련이 필요 없는 메모리 프레임워크인 WorldTrace를 제안합니다. WorldTrace는 각 요약 슬롯에 고유한, 훈련 데이터 분포 내의 가상 위치를 할당하여 압축된 메모리를 어드레서블하게 유지합니다. 이 어드레서블 캐시 내에서, 우리는 두 가지 메모리 압축 방식을 연구했습니다. WorldTrace-Field는 시간적 일관성을 위해 과거 정보를 압축하고, WorldTrace-Landmark는 감지된 전환 지점에서 장면의 정확한 내용을 저장하여 에피소드 기억을 가능하게 합니다. 또한, 긴 우회로를 거친 후에도 이전 방문 장면을 재구성할 수 있는지 평가하는 벤치마크인 LoopBench를 소개합니다. 실험 결과, WorldTrace-Field는 시간적 일관성을 +15.5% 향상시키고, WorldTrace-Landmark는 에피소드 기억을 +19.5% 향상시키는 것으로 나타났으며, 이는 재훈련 없이 시각적으로 지속적인 생성이 가능하도록 합니다.

Original Abstract

We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then fall outside the range seen during training and the model struggles to retrieve the relevant visual information through attention. Moreover, naively compressing the cache in the RoPE-rotated space corrupts memory by averaging together incompatible positional phases. To address this, we propose WorldTrace, a training-free memory framework for long-horizon visual persistence. WorldTrace keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position. Within this addressable cache, we study two memory compression approaches: WorldTrace-Field compresses history for temporal coherence, while WorldTrace-Landmark stores verbatim scene traces at detected transitions for episodic recall. We further introduce LoopBench, a benchmark evaluating whether a compressed cache can reconstruct a previously visited scene after a long detour. WorldTrace-Field improves temporal consistency by +15.5%, and WorldTrace-Landmark improves episodic recall by +19.5% on LoopBench, extending visually persistent generation without retraining.

0 Citations
0 Influential
12 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!