2606.31232v1 Jun 30, 2026 cs.AI

델타-JEPA: 잠재적 차이 디코딩을 통한 동작 민감적인 세계 모델 학습

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

Hongzhu Yi
Hongzhu Yi
Citations: 9
h-index: 2
Zhenghao Zhang
Zhenghao Zhang
Citations: 12
h-index: 2
Tianyu Zong
Tianyu Zong
Citations: 153
h-index: 5
Yuanxiang Wang
Yuanxiang Wang
Citations: 10
h-index: 2
Tao Yu
Tao Yu
Citations: 8
h-index: 2
Yujia Yang
Yujia Yang
Citations: 7
h-index: 2
Zhenyu Guan
Zhenyu Guan
Citations: 4
h-index: 1
Jungang Xu
Jungang Xu
Citations: 11
h-index: 2
Bingkang Shi
Bingkang Shi
Citations: 43
h-index: 3
Xing Chen
Xing Chen
Citations: 250
h-index: 5
Tiankun Yang
Tiankun Yang
Citations: 7
h-index: 1
Guoqing Chao
Guoqing Chao
Citations: 22
h-index: 2
Chenxi Bao
Chenxi Bao
Citations: 0
h-index: 0
Jingjing Zhou
Jingjing Zhou
Citations: 0
h-index: 0

계획을 위한 시각적 세계 모델 학습에는 작용에 민감하게 반응하는 간결한 잠재적 동역학이 필요하지만, 재구성 없이 결합된 임베딩 목표는 작용에 둔감한 표현으로 이어질 수 있습니다. 본 논문에서는 잠재적 순방향 예측에 잠재적 차이 동작 디코더(LDAD)를 추가하여 엔드투엔드 방식으로 재구성을 하지 않는 세계 모델인 Delta-JEPA를 제안합니다. 기존의 역 디코더가 연결된 종단 임베딩에서 작용을 추론하는 것과 달리, LDAD는 연속적인 관측값 사이의 잠재적 변위로부터 실행된 작용을 재구성합니다. 이러한 변위 수준의 감독은 전이 기하학을 직접적으로 규제합니다. 인접한 임베딩은 작용 정보를 잃지 않고서는 붕괴될 수 없으며, 다양한 작용은 계획 기반 추론을 위해 구별 가능한 잠재적 변화를 유도하도록 장려됩니다. Delta-JEPA는 잠재적 예측과 작용 재구성만 사용하며, 픽셀 재구성과 분포 매칭 정규화 방법을 사용하지 않습니다. 네 가지 시각적 연속 제어 작업에서 Delta-JEPA는 JEPA 기반 및 표현 학습 세계 모델 기준 성능을 능가하는 결과를 보였습니다. 분석 결과, 변위 기반의 작용 디코딩이 종단 연결보다 일관되게 효과적인 것으로 나타났으며, 작용 민감성 분석은 더 명확한 작용 조건화된 잠재적 반응을 보여주었습니다. 이러한 결과는 잠재적 차이에 대한 감독이 붕괴 방지 및 작용 민감성을 갖춘 세계 모델 학습을 위한 간단하고 효과적인 메커니즘임을 시사합니다.

Original Abstract

Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations. We propose Delta-JEPA, an end-to-end reconstruction-free world model that augments latent forward prediction with a Latent Difference Action Decoder (LDAD). Unlike inverse decoders that infer actions from concatenated endpoint embeddings, LDAD reconstructs the executed action from the latent displacement between consecutive observations. This displacement-level supervision directly regularizes transition geometry: adjacent embeddings cannot collapse without losing action information, and different actions are encouraged to induce distinguishable latent changes for rollout-based planning. Delta-JEPA uses only latent prediction and action reconstruction, avoiding pixel reconstruction and distribution-matching regularizers. Across four visual continuous-control tasks, Delta-JEPA improves planning over JEPA-based and representation-learning world model baselines. Ablations show that displacement-based action decoding is consistently more effective than endpoint concatenation, and action-sensitivity analyses show clearer action-conditioned latent responses. These results indicate that supervising latent differences is a simple and effective mechanism for collapse-resistant and action-sensitive world model learning.

2 Citations
2 Influential
2.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!