비디오 모델을 네이티브 4차원 렌더러로 활용: 애니메이션 메시 기반의 세계 정보 연동
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
사전 학습된 비디오 확산 모델은 원하는 장면 상태가 애니메이션 메시, 카메라 경로 및 참조 이미지에 의해 이미 정의되어 있을 때 렌더러 역할을 할 수 있습니다. 이러한 4차원 생성 렌더링 환경에서 중요한 질문은 다음과 같습니다: 어떤 이미지 형식의 조건이 비디오 기반 모델이 카메라 움직임과 장면 내 애니메이션 모두를 만족하도록 만들 수 있는가? 우리는 DAR(DAR)라는 참조 기반 렌더러를 제안합니다. DAR은 Plücker 광선을 사용하여 Wan2.2 카메라 제어를 확장하여, 대신 카메라와 지오메트리를 결합한 인터페이스를 사용합니다. DAR은 애니메이션 메시에서 파생된 신경망 기반의 4차원 G-버퍼(트래킹 정보, 월드 위치 및 노멀)를 생성하고, 사전 학습된 이미지-비디오 관계를 유지하면서 확장된 제어 어댑터를 통해 이 정보를 주입합니다. 핵심 설계 선택은 트래킹 정보와 월드 위치 쌍입니다. 트래킹 정보는 텍스처가 적용되어야 할 지속적인 표면 요소를 식별하며, 월드 위치는 해당 요소의 현재 장면 좌표 상태를 나타냅니다. 노멀은 로컬 형상을 제공합니다. 깊이 정보와 보정된 광선을 사용하면 원칙적으로 3차원 정보를 복구할 수 있지만, 깊이는 카메라 의존적인 값이며, 여기에는 카메라 움직임과 객체 움직임이 혼합되어 있습니다. DAR-4D 벤치마크에서 LoRA DAR은 PSNR 23.22, SSIM 0.895 및 LPIPS 0.134를 달성했으며, 이는 기존 Wan2.2-Depth 모델보다 PSNR 기준 1.54dB 향상된 결과입니다. 전체 파인튜닝을 통해 PSNR은 25.36까지, SSIM은 0.917까지 향상되었습니다. 추가 실험 결과, 월드 위치를 깊이 정보로 대체하면 모든 체크포인트에서 PSNR이 1.26~1.55dB 감소하는 것으로 나타났습니다. 이는 트래킹 정보와 월드 위치의 조합이 실용적인 4차원 렌더링 조건임을 뒷받침합니다.
Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D generative rendering setting raises a representation question: what image-format condition lets a video backbone obey both camera motion and scene-internal animation? We propose DAR, a reference-guided renderer that extends Wan2.2 camera control from Plücker rays alone to a joint camera-plus-geometry interface. DAR projects a neural 4D G-buffer (tracking, world position, and normal) from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior. The central design choice is the pair of tracking and world position. Tracking identifies the persistent surface element that should carry appearance; world position gives its current scene-coordinate state; normal supplies local shape. Depth plus calibrated rays can recover 3D in principle, but depth is a camera-dependent chart in which camera and object motion are mixed. On the 68-case DAR-4D benchmark, LoRA DAR reaches PSNR 23.22, SSIM 0.895, and LPIPS 0.134, improving over off-the-shelf Wan2.2-Depth by 1.54 dB PSNR; a full fine-tune reaches PSNR 25.36 and SSIM 0.917. Matched ablations show that replacing world position by depth reduces PSNR by 1.26--1.55 dB at every checkpoint, supporting tracking+world-position correspondence as a practical 4D rendering condition.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.