2606.20104v1 Jun 18, 2026 cs.LG

감각-운동 세계 모델: 역동학을 통한 행동 기반의 인지

Sensorimotor World Models: Perception for Action via Inverse Dynamics

Bernhard Scholkopf
Bernhard Scholkopf
Citations: 142
h-index: 6
Randall Balestriero
Randall Balestriero
Citations: 71
h-index: 3
P. Ivashkov
P. Ivashkov
Citations: 41
h-index: 4

행동 기반의 인지는 세상에 대한 표현이 시각적 정확성뿐만 아니라, 행동과의 관련성에 의해 형성되어야 함을 시사합니다. 동시에, 잠재적인 JEPA 스타일의 세계 모델은 고차원 데이터로부터 간결한 예측 상태를 학습하여 미래 상태를 예측하는 것을 목표로 하지만, 이러한 모델의 end-to-end 훈련은 표현이 붕괴될 수 있기 때문에 쉽지 않습니다. 본 논문에서는 역동학 정규화를 활용하여 end-to-end 방식으로 학습되는 감각-운동 세계 모델(SMWM)을 제안합니다. 이 단일 정규화 기법은 표현 붕괴를 방지하고, 동시에 행동과 관련된 표현을 유도하는 역할을 합니다. SMWM은 상태 전이에 대한 정보를 보존하도록 잠재 변수를 강제하여, 환경의 제어 가능한 요소에 대한 모델의 편향성을 높이고, 제어 불가능한 잡음을 제거합니다. 결과적으로, 본 연구는 frozen encoder, 지수 이동 평균, 또는 복잡한 잠재 변수 정규화 없이, 오프라인 데이터와 보상 신호 없이도 안정적인 잠재 세계 모델을 학습할 수 있음을 보여줍니다. 실험적으로, SMWM은 간결하고 해석 가능한 잠재 공간을 학습하며, 간단한 2D 및 3D 제어 작업에서 경쟁력 있는 계획 성능을 달성합니다.

Original Abstract

Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions. At the same time, latent JEPA-style world models advocate learning compact predictive states from high-dimensional observations to facilitate the prediction of future states, but end-to-end training of these models is nontrivial because representations may collapse if our only goal is to construct a latent state that is easy to predict. We introduce a sensorimotor world model (SMWM): a latent world model trained end-to-end with inverse dynamics regularization. This single regularizer addresses both issues: it prevents representation collapse and induces action-aligned representations. By forcing latent states to preserve information about the action underlying a transition, it biases the model toward the controllable degrees of freedom of the environment while discarding uncontrollable distractors. This yields stable latent world models trained from offline, reward-free trajectories, without frozen encoders, exponential moving averages, or complex latent regularizers. Empirically, SMWM learns compact, interpretable latent spaces and enables competitive planning performance across simple 2D and 3D control tasks.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!