2601.16296v1 Jan 22, 2026 cs.CV

Memory-V2V: 메모리를 활용한 비디오-투-비디오 확산 모델

Memory-V2V: Augmenting Video-to-Video Diffusion Models with Memory

Dohun Lee
Dohun Lee
Citations: 37
h-index: 4
C. Huang
C. Huang
Citations: 155
h-index: 6
Xuelin Chen
Xuelin Chen
Citations: 8
h-index: 1
Duygu Ceylan
Duygu Ceylan
Citations: 178
h-index: 7
Hyeonho Jeong
Hyeonho Jeong
Citations: 302
h-index: 6
J. Ye
J. Ye
Citations: 257
h-index: 10

최근 개발된 비디오-투-비디오 확산 모델은 사용자가 제공한 비디오의 외관, 움직임 또는 카메라 움직임을 수정하여 편집하는 데 놀라운 성과를 거두었습니다. 그러나 실제 비디오 편집은 종종 반복적인 과정이며, 사용자는 여러 단계의 상호 작용을 통해 결과를 개선합니다. 이러한 다단계 설정에서, 현재의 비디오 편집기는 순차적인 편집 과정에서 일관성을 유지하는 데 어려움을 겪습니다. 본 연구에서는 다단계 비디오 편집에서 발생하는 일관성 문제를 처음으로 다루고, 명시적인 메모리를 활용하여 기존의 비디오-투-비디오 모델을 향상시키는 간단하면서도 효과적인 프레임워크인 Memory-V2V를 소개합니다. Memory-V2V는 이전에 편집된 비디오를 저장한 외부 캐시를 활용하여 정확한 검색 및 동적 토큰화 전략을 통해 현재 편집 단계에서 이전 결과를 반영합니다. 또한, 중복을 줄이고 계산 오버헤드를 완화하기 위해 DiT 백본 내에 학습 가능한 토큰 압축기를 제안하여 불필요한 조건 토큰을 압축하면서 필수적인 시각적 정보를 유지하고, 전체적으로 30%의 속도 향상을 달성했습니다. Memory-V2V는 비디오 신규 뷰 합성 및 텍스트 기반의 긴 비디오 편집과 같은 어려운 작업에서 성능을 검증했습니다. 광범위한 실험 결과, Memory-V2V는 최소한의 계산 오버헤드로 비디오 간의 일관성이 훨씬 뛰어나며, 동시에 기존 최고 성능 모델을 능가하거나 동등한 수준의 특정 작업 성능을 유지하는 것을 확인했습니다. 프로젝트 페이지: https://dohunlee1.github.io/MemoryV2V

Original Abstract

Recent foundational video-to-video diffusion models have achieved impressive results in editing user provided videos by modifying appearance, motion, or camera movement. However, real-world video editing is often an iterative process, where users refine results across multiple rounds of interaction. In this multi-turn setting, current video editors struggle to maintain cross-consistency across sequential edits. In this work, we tackle, for the first time, the problem of cross-consistency in multi-turn video editing and introduce Memory-V2V, a simple, yet effective framework that augments existing video-to-video models with explicit memory. Given an external cache of previously edited videos, Memory-V2V employs accurate retrieval and dynamic tokenization strategies to condition the current editing step on prior results. To further mitigate redundancy and computational overhead, we propose a learnable token compressor within the DiT backbone that compresses redundant conditioning tokens while preserving essential visual cues, achieving an overall speedup of 30%. We validate Memory-V2V on challenging tasks including video novel view synthesis and text-conditioned long video editing. Extensive experiments show that Memory-V2V produces videos that are significantly more cross-consistent with minimal computational overhead, while maintaining or even improving task-specific performance over state-of-the-art baselines. Project page: https://dohunlee1.github.io/MemoryV2V

1 Citations
0 Influential
5 Altmetric
26.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!