델타-JEPA: 잠재적 차이 디코딩을 통한 동작 민감적인 세계 모델 학습
Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding
계획을 위한 시각적 세계 모델 학습에는 작용에 민감하게 반응하는 간결한 잠재적 동역학이 필요하지만, 재구성 없이 결합된 임베딩 목표는 작용에 둔감한 표현으로 이어질 수 있습니다. 본 논문에서는 잠재적 순방향 예측에 잠재적 차이 동작 디코더(LDAD)를 추가하여 엔드투엔드 방식으로 재구성을 하지 않는 세계 모델인 Delta-JEPA를 제안합니다. 기존의 역 디코더가 연결된 종단 임베딩에서 작용을 추론하는 것과 달리, LDAD는 연속적인 관측값 사이의 잠재적 변위로부터 실행된 작용을 재구성합니다. 이러한 변위 수준의 감독은 전이 기하학을 직접적으로 규제합니다. 인접한 임베딩은 작용 정보를 잃지 않고서는 붕괴될 수 없으며, 다양한 작용은 계획 기반 추론을 위해 구별 가능한 잠재적 변화를 유도하도록 장려됩니다. Delta-JEPA는 잠재적 예측과 작용 재구성만 사용하며, 픽셀 재구성과 분포 매칭 정규화 방법을 사용하지 않습니다. 네 가지 시각적 연속 제어 작업에서 Delta-JEPA는 JEPA 기반 및 표현 학습 세계 모델 기준 성능을 능가하는 결과를 보였습니다. 분석 결과, 변위 기반의 작용 디코딩이 종단 연결보다 일관되게 효과적인 것으로 나타났으며, 작용 민감성 분석은 더 명확한 작용 조건화된 잠재적 반응을 보여주었습니다. 이러한 결과는 잠재적 차이에 대한 감독이 붕괴 방지 및 작용 민감성을 갖춘 세계 모델 학습을 위한 간단하고 효과적인 메커니즘임을 시사합니다.
Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations. We propose Delta-JEPA, an end-to-end reconstruction-free world model that augments latent forward prediction with a Latent Difference Action Decoder (LDAD). Unlike inverse decoders that infer actions from concatenated endpoint embeddings, LDAD reconstructs the executed action from the latent displacement between consecutive observations. This displacement-level supervision directly regularizes transition geometry: adjacent embeddings cannot collapse without losing action information, and different actions are encouraged to induce distinguishable latent changes for rollout-based planning. Delta-JEPA uses only latent prediction and action reconstruction, avoiding pixel reconstruction and distribution-matching regularizers. Across four visual continuous-control tasks, Delta-JEPA improves planning over JEPA-based and representation-learning world model baselines. Ablations show that displacement-based action decoding is consistently more effective than endpoint concatenation, and action-sensitivity analyses show clearer action-conditioned latent responses. These results indicate that supervising latent differences is a simple and effective mechanism for collapse-resistant and action-sensitive world model learning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.