2608.07409v1 Aug 07, 2026 cs.CV

UniJEPA: 작업에 독립적인 시각적 세계 모델링을 위한 통합된 공동 임베딩 예측 아키텍처

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

Haoran Xu
Haoran Xu
Citations: 54
h-index: 5
Dawei Liu
Dawei Liu
Citations: 43
h-index: 2
Andriana Lanji
Andriana Lanji
Citations: 0
h-index: 0
Jin Li
Jin Li
Citations: 0
h-index: 0
Mei Chen
Mei Chen
Citations: 0
h-index: 0
Yuying Tian
Yuying Tian
Citations: 0
h-index: 0

공동 임베딩 예측 아키텍처(JEPAs)는 압축된 잠재 공간에서 세계 모델의 자기 지도 학습을 위한 체계적인 프레임워크로 부상했지만, 기존 방법들은 단편적입니다. 일부는 잠재 공간 내에서 이미지의 마스크 처리된 부분을 예측합니다 (I-JEPA), 다른 일부는 전역 광도학적 변환을 예측하는 것을 학습합니다 (Image World Models), 그리고 비디오 수준의 JEPAs는 미래 시간 상태를 예측하고 동작에 조건화된 계획을 위해 추가적으로 훈련됩니다 (V-JEPA~2, DINO-World, DINO-WM). 이러한 목표들은 별개의 인코더, 예측기 및 안티 콜랩스 정규화를 갖춘 개별적인 방식으로 취급되어, 단일 모델이 이미지 수준과 비디오 수준의 세계 모델링을 통합하는 것을 어렵게 만듭니다. 본 논문에서는 UniJEPA를 제안합니다. UniJEPA는 하나의 공유된 잠재 공간에서 광도학적 예측 (이미지 수준 변환)과 시간 예측 (비디오 수준 다음 상태 동역학)을 동시에 학습하는 통합된 JEPA입니다. 차세대 임베딩 예측 손실과 가우시안 정규화로 구성된 단일 엔드 투 엔드 목표는 EMA, 스톱-그라디언트 또는 사전 훈련된 인코더 없이 원본 픽셀로부터 학습 가능한 안티 콜랩스 인코더-예측기 쌍을 제공합니다. 동일한 잠재 공간은 제어 가능한 추상화를 지원하며, 광도학적 예측은 불변 구조를 학습하고 시간 예측은 동등 변환 동역학을 학습합니다. 오프라인 트랙터리에 대한 동작 조건화 후 훈련을 통해 UniJEPA는 목표 특징을 예측 대상으로 사용하여 제로-샷 계획 기능을 제공합니다. 이미지, 비디오 및 제어 벤치마크에서 UniJEPA는 특정 작업에 특화된 JEPAs와 동등하거나 뛰어난 성능을 보이며, 단일 손실 하이퍼파라미터만 필요하고, 유사한 정확도로 생성적 세계 모델보다 최대 수십 배 빠르게 계획할 수 있습니다.

Original Abstract

Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods are fragmented: some predict masked parts of a single image in latent space (I-JEPA), others learn to predict global photometric transformations (Image World Models), while video-scale JEPAs predict future temporal states and are post-trained for action-conditioned planning (V-JEPA~2, DINO-World, DINO-WM). These objectives are treated as distinct recipes with separate encoders, predictors, and anti-collapse regularizers, hindering a single model from unifying image-level and video-level world modeling. We present UniJEPA, a unified JEPA that jointly learns photometric prediction (image-level transformations) and temporal prediction (video-level next-state dynamics) in one shared latent space. A single end-to-end objective, composed of a next-embedding prediction loss and a Gaussian regularizer, yields a provably anti-collapse encoder-predictor pair trainable from raw pixels without EMA, stop-gradient, or pre-trained encoders. We show that the same latent space supports controllable abstraction: photometric prediction learns invariant structure while temporal prediction learns equivariant dynamics. After action-conditioned post-training on offline trajectories, UniJEPA enables zero-shot planning by treating goal features as prediction targets. On image, video, and control benchmarks, UniJEPA matches or surpasses task-specific JEPAs while requiring a single loss hyperparameter, and plans up to tens of times faster than generative world models at comparable accuracy.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!