렌더링, 디코딩하지 않기: 잠재 구조 분리 기반 가중치 공간 세계 모델
Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement
방대한 양의 비표시된 동영상 데이터를 활용하여 세계 모델을 훈련하는 것은 완전 자율 지능으로 나아가기 위한 중요한 단계입니다. 그러나 현재의 주요 패러다임은 원시 픽셀을 불투명한 잠재 공간으로 인코딩하고 복원을 위해 복잡한 디코더에 의존하기 때문에 이러한 모델은 계산 비용이 많이 들고 해석하기 어렵습니다. 우리는 이 문제를 해결하기 위해 NOVA라는 세계 모델링 프레임워크를 소개합니다. NOVA는 시스템 상태를 보조적인 좌표 기반 암시적 신경 표현(INR)의 가중치와 편향으로 표현합니다. 이 구조화된 표현은 분석적으로 렌더링되므로 디코더 병목 현상을 제거하는 동시에 압축성, 이식성 및 제로샷 초해상도를 제공합니다. 또한, 대부분의 잠재 행동 모델과 마찬가지로 NOVA는 행동 매칭 목표를 통해 컨텍스트에 의존적인 동영상 생성기로 증류될 수 있습니다. 놀랍게도, 추가적인 손실 함수나 적대적 목표를 사용하지 않고도 NOVA는 배경, 전경 및 프레임 간 움직임과 같은 구조적 장면 구성 요소를 분리할 수 있습니다. 이를 통해 사용자는 콘텐츠 또는 동력학 중 하나를 변경하면서 다른 하나를 손상시키지 않고 편집할 수 있습니다. 우리는 여러 가지 어려운 데이터 세트에서 우리의 프레임워크를 검증하여 강력한 제어 가능 예측 성능을 달성했으며, 단일 소비자 GPU에서 약 40M개의 매개변수로 작동합니다. 궁극적으로, INR과 같은 구조화된 표현은 잠재적인 동역학에 대한 우리의 이해를 높일 뿐만 아니라 몰입감 있고 맞춤화 가능한 가상 경험을 위한 길을 열어줍니다.
Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing paradigm of encoding raw pixels into opaque latent spaces and relying on heavy decoders for reconstruction leaves these models computationally expensive and uninterpretable. We address this problem by introducing NOVA, a world modelling framework that represents the system state as the weights and biases of an auxiliary coordinate-based implicit neural representation (INR). This structured representation is analytically rendered, which eliminates the decoder bottleneck while conferring compactness, portability, and zero-shot super-resolution. Furthermore, like most latent action models, NOVA can be distilled into a context-dependent video generator via an action-matching objective. Surprisingly, without resorting to auxiliary losses or adversarial objectives, NOVA can disentangle structural scene components such as background, foreground, and inter-frame motion, enabling users to edit either content or dynamics without compromising the other. We validate our framework on several challenging datasets, achieving strong controllable forecasting while operating on a single consumer GPU at $\sim$40M parameters. Ultimately, structured representations like INRs not only enhance our understanding of latent dynamics but also pave the way for immersive and customisable virtual experiences.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.