2608.05720v1 Aug 06, 2026 cs.CV

PhyLatent: JEPA 세계 모델을 위한 역학적 특성을 반영하는 표현 학습

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Haojie Ren
Haojie Ren
Citations: 43
h-index: 3
Ziying Song
Ziying Song
Citations: 23
h-index: 3

본 논문에서는 Joint Embedding Predictive Architecture (JEPA) 기반의 세계 모델에 대한 역학적 특성을 고려한 학습 목표인 PhyLatent를 제안합니다. 핵심적인 관찰 결과는 전역 잠재 변수의 붕괴를 막는 것만으로는 표현이 물리적 상태와 행동 결과를 보존하는 것을 보장할 수 없다는 것입니다. 우리는 JEPA 세계 모델에서 발생하는 세 가지 문제점을 지적했습니다: 물리적 불변성 붕괴, 물리적 식별 가능성 붕괴 및 반사실적 역학 붕괴입니다. PhyLatent는 물리적 상태 기반 학습, 미래 표현 정렬, 정적 시각적 불변성, 반사실적 분기 분리 및 잠재 변수 노이즈 제거를 통해 세 가지 학습 경로를 제공하여 이러한 문제점을 해결합니다. OGBench-Cube 데이터셋에서 PhyLatent는 세 가지 문제점 발생률을 각각 15.60%, 6.71%, 8.41%에서 7.53%, 0.95%, 4.62%로 감소시키고, 모델 예측 제어 (MPC) 성공률을 70.0%에서 78.1%로 향상시켰습니다. 동일한 구조와 계획기를 사용하여 TwoRooms 데이터셋에서도 성공률을 81.0%에서 98.0%로 더욱 향상시켰으며, Reacher 및 PushT 데이터셋에서도 경쟁력 있는 성능을 유지했습니다. 이러한 결과는 전역적인 붕괴 방지만으로는 신뢰할 수 있는 JEPA 세계 모델 상태 공간을 학습하는 데 충분하지 않다는 것을 보여줍니다.

Original Abstract

We propose PhyLatent, a dynamics-relevant training objective for JointEmbedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not ensure that a representation preserves physical states and action consequences. We identify three failure modes in JEPA world models: physical invariance collapse, physical identifiability collapse, and counterfactual dynamics collapse. PhyLatent addresses them through three training pathways: physical invariance, physical identifiability, and counterfactual dynamics, implemented with physical state grounding, future representation alignment, static visual invariance, counterfactual branch separation, and latent denoising. On OGBench-Cube, PhyLatent reduces the three failure rates from 15.60%, 6.71%, and 8.41% to 7.53%, 0.95%, and 4.62%, respectively, and improves model predictive control (MPC) success from 70.0% to 78.1%. With the same architecture and planner, it further improves success from 81.0% to 98.0% on TwoRooms and remains competitive on Reacher and PushT. These results show that global non-collapse alone is insufficient for learning a reliable JEPA worldmodel state space.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!