2608.08553v1 Aug 09, 2026 cs.CV

MotionCraft: 희소 어텐션을 활용한 잠재 공간 모델 기반 시각적 업스케일링

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

Rong Fu
Rong Fu
Citations: 45
h-index: 4
Simon Fong
Simon Fong
Citations: 26
h-index: 2
Chunlei Meng
Chunlei Meng
Citations: 4
h-index: 2
Shuo Yin
Shuo Yin
Citations: 16
h-index: 3
Wangyu Wu
Wangyu Wu
Citations: 12
h-index: 2
Yongtai Liu
Yongtai Liu
Citations: 1
h-index: 1
Yangcheng Zeng
Yangcheng Zeng
Citations: 1
h-index: 1
Zijian Zhang
Zijian Zhang
Citations: 28
h-index: 2
Xiaowen Ma
Xiaowen Ma
Citations: 733
h-index: 16
Yingrui Ji
Yingrui Ji
Citations: 20
h-index: 3
Sicheng Li
Sicheng Li
Citations: 10
h-index: 1
Chenhao Wang
Chenhao Wang
Citations: 0
h-index: 0

비디오 초해상도(VSR)는 저해상도 비디오 입력을 통해 고화질 비디오를 복원하는 기술로, 모바일 촬영부터 스트리밍 및 보존 복원에 이르기까지 다양한 분야에서 핵심적인 역할을 합니다. 기존 방법들은 지역 디테일의 정확성, 장거리 시공간 모델링, 인지적 현실감, 효율성 간의 균형을 맞추는 데 어려움을 겪습니다. 컨볼루션 기반 정렬 기법은 지역 구조를 보존하지만, 움직임이 크거나 왜곡이 복잡한 경우 성능 저하가 발생합니다. 트랜스포머 기반 방법은 장거리 의존성을 포착할 수 있지만, 계산 가능하도록 아키텍처 또는 알고리즘을 수정해야 합니다. 최근의 잠재 공간 또는 확산 모델 기반 생성기는 풍부한 텍스처를 합성하지만, 일관성을 유지하기 위해 특수한 시간 제약 조건이 필요합니다. 본 논문에서는 세계 모델에서 영감을 받아 움직임 인지 잠재 상태 예측을 통해 복원을 수행하고, 명시적인 사용자 인터페이스를 갖춘 적응형 희소 어텐션을 통합한 VSR 프레임워크인 MotionCraft를 제시합니다. MotionCraft는 강력한 움직임 융합, 지역성과 대상 비지역 상호 작용의 균형을 맞추는 잠재 공간 트랜스포머, 그리고 시간적 일관성이 뛰어나고 고품질의 복원을 실시간으로 제공할 수 있는 경량 조건부 디코더를 결합합니다. 실험 결과 MotionCraft는 뛰어난 복원 및 인지 성능을 달성하며, 시간적 부드러움과 복원 정확성 간의 예측 가능한 균형을 제공하는 것을 확인했습니다.

Original Abstract

Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!