ShadowDancer: 비디오 및 그림자를 활용하여 통일된 동역학 표현을 학습함으로써, 비디오 월드 모델에 어떤 동작이라도 가르치는 방법
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
본 논문에서는 ShadowDancer라는 새로운 접근 방식을 제시합니다. 이는 상호작용 가능한 비디오 월드 모델의 모든 동작과 프레임 단위 제어를 가능하게 합니다. 기존 방식들은 표현 방식에 어려움을 가지고 있는데, 일부는 동작을 느슨하게 인코딩하여 모델이 임의로 구현하도록 하지만, 다른 일부는 정교한 신호를 사용하여 특정 유형의 동작만 처리하며 다양한 동역학적 환경에서 정확한 제어를 어렵게 만듭니다. 데모 비디오는 프레임별로 모든 동역학을 명시하는 자연스러운 해결책이지만, 비디오는 특정 외형만을 보여주며 이는 기본 동역학의 '그림자'에 불과하기 때문에, 데모를 통해 학습된 동작은 새로운 환경으로 잘 전달되지 않습니다. ShadowDancer는 두 가지 핵심적인 혁신을 통해 이러한 문제를 해결합니다. (1) shadow pairs(그림자 쌍): 동일한 동역학을 독립적으로 재샘플링된 외형 하에서 반복하는 비디오 쌍으로, Shadow Library를 통해 대규모로 생성되며, 이를 통해 특정 동역학 패밀리에 대한 정확한 제어가 가능해집니다. (2) cross-shadow prediction(크로스 그림자 예측): 한 그림자를 다른 그림자로부터 예측하여 동작을 학습하며, 이 과정에서 페어링이 재샘플링하는 내용은 무시되고 유지되는 내용만이 동작으로 남게 되어 통일된 동역학 표현을 생성하고 이를 통해 블록-원인 관계를 갖는 월드 모델을 구동합니다. 따라서 어떤 데모 클립도 재사용 가능한 동작 자산이 되며, 별도의 동작 레이블, 모션 추정기 또는 미세 조정 없이 새로운 환경에서 재생될 수 있습니다. 실험 결과는 다양한 동역학 패밀리에 걸쳐 기존의 강력한 잠재 행동 및 상호작용 월드 모델에 비해 향상된 동작 전이 및 긴 동작 실행 성능을 보여주며, 평균적으로 맹검 테스트에서 86%의 우수한 결과를 보였습니다. 비디오 시연은 https://ShadowDancer-1.github.io 에서 확인할 수 있습니다.
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.