AnchorEdit: 인과적 기억을 활용한 다중 단계 이미지 편집에서 시간적 일관성 유지
AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory
다중 단계 이미지 편집은 반복적인 디자인 과정에 필수적이지만, 기존 모델들은 종종 연속적인 단계에서 객체의 동일성 변화와 오류 누적 문제를 겪습니다. 기존 연구에서는 비디오 사전 지식을 활용하여 일관성을 확보하려 하지만, 이러한 접근 방식은 양방향 주의 메커니즘에 의존하며 이는 상호 작용 기반 편집의 인과적이고 순차적인 특성과 근본적으로 일치하지 않습니다. 본 논문에서는 고해상도 및 장기 다중 단계 편집을 위해 특별히 설계된 최초의 자기 회귀(AR) 확산 모델 프레임워크인 AnchorEdit을 제안합니다. AnchorEdit은 비디오 사전 지식과 인과적 추론 사이의 간극을 세 가지 단계의 학습 과정을 통해 좁힙니다: 객체 동일성을 유지하는 단일 단계 사전 학습, 새로운 자기 회귀 강제 미세 조정(exposure bias를 완화하기 위한 self-rollout 전략 활용), 효율적인 4단계 생성을 위한 일관성 증류. 추론 과정에서, 초기 객체의 동일성을 고정하고 확장된 편집 경로 전체에 걸쳐 안정적인 결과를 보장하기 위해 메모리 메커니즘을 도입합니다. 성능 평가를 위해, 장기적인 안정성을 검증하도록 설계된 새로운 고해상도 다중 단계 편집 벤치마크를 제공합니다. 광범위한 실험 결과는 AnchorEdit이 최첨단 결과를 달성하며, 10번 이상의 상호 작용에서도 뛰어난 객체 충실도와 지시 사항 준수를 유지한다는 것을 보여줍니다.
Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successive steps. While existing research leverages video priors for consistency, their reliance on bidirectional attention is fundamentally misaligned with the causal, sequential nature of interactive editing. In this paper, we propose AnchorEdit, the first autoregressive (AR) diffusion-based framework designed specifically for high-resolution, long-term multi-turn editing. AnchorEdit bridges the gap between video priors and causal inference through a three-stage training curriculum: identity-preserving sing-turn pretraining, causal AR forcing fine-tuning with a novel self-rollout strategy to mitigate exposure bias, and consistency distillation for efficient 4-step generation. During inference, we introduce a memory mechanism to anchor the initial subject identity and ensure stable extrapolation across extended editing trajectories. To evaluate performance, we provide a new high-resolution multi-turn editing benchmark designed to stress-test long-horizon stability. Extensive experiments demonstrate that AnchorEdit achieves state-of-the-art results, maintaining exceptional subject fidelity and instruction following even over 10+ interaction rounds.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.