2607.21529v1 Jul 23, 2026 cs.CV

ElasticTTT: 사전 정보 보존을 위한 테스트 시간 조정 방식 - 비디오 편집

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Miao Liu
Miao Liu
Citations: 78
h-index: 5
Yueyi Liu
Yueyi Liu
Citations: 0
h-index: 0
Chi Zhang
Chi Zhang
Citations: 34
h-index: 1
Sen Cui
Sen Cui
Citations: 0
h-index: 0

사전 학습된 확산 모델에 대한 테스트 시간 조정(TTT)은 비디오 편집 분야에서 강력한 패러다임으로 부상했습니다. 그러나 생성 모델의 분포 매핑 특성과 표준 TTT의 단일 지점 최적화 간에는 근본적인 불일치가 존재합니다. 본 논문에서는 이러한 불일치가 *사전 정보 붕괴(Prior Collapse)*를 유발한다는 것을 보여줍니다. 이는 모델이 텍스트 조건과 공간 잠재 변수를 무시하여 생성 결과가 원본 비디오로 수렴하거나, 서로 다른 영역의 특징을 혼합시키는 현상입니다. 이러한 문제를 해결하기 위해, 본 논문에서는 사전 분포를 보존하고 생성적 유연성을 회복하는 새로운 프레임워크인 **ElasticTTT**를 제안합니다. 구체적으로, 급격한 기억 최소값을 방지하기 위한 *타겟 분포 정규화(Target Distribution Regularization)*, 원본 편향에서 벗어나 추론을 안내하기 위한 *대조적 CFG(Contrastive CFG)*, 그리고 편집되지 않은 영역을 보존하기 위한 *비동기 노이즈 스케줄(Asynchronous Noise Schedule)*을 제안합니다. 이론적 분석으로 뒷받침된 광범위한 실험 결과는 ElasticTTT가 기본 모델의 생성적 사전 정보를 성공적으로 보존하며, 단일 예제 기반 비디오 편집에서 최첨단 성능을 달성한다는 것을 보여줍니다.

Original Abstract

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.

1 Citations
0 Influential
2.5 Altmetric
13.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!