MotionCraft: 희소 어텐션을 활용한 잠재 공간 모델 기반 시각적 업스케일링
MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling
비디오 초해상도(VSR)는 저해상도 비디오 입력을 통해 고화질 비디오를 복원하는 기술로, 모바일 촬영부터 스트리밍 및 보존 복원에 이르기까지 다양한 분야에서 핵심적인 역할을 합니다. 기존 방법들은 지역 디테일의 정확성, 장거리 시공간 모델링, 인지적 현실감, 효율성 간의 균형을 맞추는 데 어려움을 겪습니다. 컨볼루션 기반 정렬 기법은 지역 구조를 보존하지만, 움직임이 크거나 왜곡이 복잡한 경우 성능 저하가 발생합니다. 트랜스포머 기반 방법은 장거리 의존성을 포착할 수 있지만, 계산 가능하도록 아키텍처 또는 알고리즘을 수정해야 합니다. 최근의 잠재 공간 또는 확산 모델 기반 생성기는 풍부한 텍스처를 합성하지만, 일관성을 유지하기 위해 특수한 시간 제약 조건이 필요합니다. 본 논문에서는 세계 모델에서 영감을 받아 움직임 인지 잠재 상태 예측을 통해 복원을 수행하고, 명시적인 사용자 인터페이스를 갖춘 적응형 희소 어텐션을 통합한 VSR 프레임워크인 MotionCraft를 제시합니다. MotionCraft는 강력한 움직임 융합, 지역성과 대상 비지역 상호 작용의 균형을 맞추는 잠재 공간 트랜스포머, 그리고 시간적 일관성이 뛰어나고 고품질의 복원을 실시간으로 제공할 수 있는 경량 조건부 디코더를 결합합니다. 실험 결과 MotionCraft는 뛰어난 복원 및 인지 성능을 달성하며, 시간적 부드러움과 복원 정확성 간의 예측 가능한 균형을 제공하는 것을 확인했습니다.
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.