2607.05238v1 Jul 06, 2026 cs.AI

MoP-JEPA: 강제 할당 예측기 혼합을 이용한 확률적 JEPA 세계 모델

MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models

Minghao Yang
Minghao Yang
Citations: 73
h-index: 4
Tianxu Lv
Tianxu Lv
Citations: 4
h-index: 1
Zhi Song
Zhi Song
Citations: 0
h-index: 0
Ximing Xing
Ximing Xing
BUAA(Beihang University), School of Software
Citations: 383
h-index: 10
Zhenchao Tang
Zhenchao Tang
Citations: 6
h-index: 1
Hanbo Huang
Hanbo Huang
Citations: 11
h-index: 2
Zhongzheng Niu
Zhongzheng Niu
Citations: 0
h-index: 0
He Bing
He Bing
Citations: 0
h-index: 0
Lusheng Wang
Lusheng Wang
Citations: 0
h-index: 0
Jianhua Yao
Jianhua Yao
Citations: 677
h-index: 11

JEPA 세계 모델은 단일의 결정론적인 예측기를 사용하여 다음 잠재 상태를 예측하며, 이 예측기는 잠재 회귀를 통해 학습됩니다. 본 연구에서는 환경이 확률적일 때 이러한 방식이 구조적으로 실패한다는 것을 보여줍니다. 분기점에서 회귀 최적화된 예측기는 후속 임베딩의 조건부 평균을 출력하는데, 이는 실제 다음 상태와는 일치하지 않는, 어떤 상태도 나타내지 못하는 지점입니다. 본 연구에서는 결정론적인 혼합 전문가 모델과 게이티드 혼합 전문가 모델 모두에서 이러한 현상이 발생함을 증명합니다. 또한 MoP-JEPA의 강제 할당 예측기는 전역 통과(single forward pass)만으로 계산 가능한, 각 후속 모드에 대해 하나의 헤드를 가진, 전환 분포의 양자화값으로 수렴한다는 것을 증명합니다. 이는 계획자가 사용하는 인터페이스입니다. 누수 없는 평가를 통해 공식 OGBench 오프라인 데이터에서 단일 예측기 기반 롤아웃을 이용한 계획은 성능이 좋지 않습니다 (성공률 $0.02$ ~ $0.09$). 반면, 본 연구에서 제안하는 예측 모드를 활용한 계획은 최대 $0.85$의 성공률을 달성하며, 모든 작업에서 결정론적 모델, 게이티드 혼합 전문가 모델, 변분 모델보다 우수한 성능을 보입니다. 다중 예측 평가는 '무임승차(coverage freeloading)'를 유발할 수 있으므로, 본 연구에서는 검증 프로토콜을 함께 제시합니다. 이는 입력에 독립적인 코드북 제어, 섞은 문맥 테스트, 라우터 게이티드 출력, 전환 정밀도 가드, 그리고 모델이 전환 그래프를 blind하게 제안하고, ground truth는 결과 확인만을 위해 사용하는 검증된 경로 기준 등으로 구성됩니다. 이 기준 하에서 본 연구의 방법은 가장 강력한 소프트웨어 기반 대안보다 모든 미로에서 훨씬 뛰어난 성능을 보입니다 ($2$ ~ $5$배). 또한 프로토콜은 해당 baseline의 낮은 점수가 실제 존재하지 않는 예측된 전환 경로를 통해 발생한다는 것을 밝혀냅니다. 동일한 모델은 실제 환경에서도 실행되어, 가장 어려운 미로에서 발표된 OGBench baseline 7개 중 두 번째로 높은 성능을 보입니다. 다중 모드 역학은 JEPA 세계 모델이 계획을 수행할 수 있는지 여부를 결정하며, 강제 할당 예측기 혼합은 최소한의 수정이며 검증 가능한 해결책입니다.

Original Abstract

JEPA world models predict the next latent state with a single deterministic predictor trained by latent regression. We show that this fails structurally when the environment is stochastic: at a branching transition, the regression-optimal predictor outputs the conditional mean of the successor embeddings, a point between the true next states that corresponds to no state at all. We prove this collapse for deterministic and gated mixture-of-experts predictors, and prove that MoP-JEPA's hard-assigned predictors converge instead to a quantizer of the transition distribution: one head per successor mode, enumerable in a single forward pass, which is the interface a planner consumes. On official OGBench offline data with leak-free evaluation, planning over single-predictor rollouts performs poorly ($0.02$--$0.09$ success) while planning over our predicted modes reaches up to $0.85$, ahead of deterministic, gated-MoE, and variational predictors on every task. Because multi-prediction evaluation invites coverage freeloading, a verification protocol is part of the method: an input-agnostic codebook control, a shuffled-context test, router-gated readouts, transition-precision guards, and a verified-route criterion in which the model proposes its transition graph blind and ground truth is used only to check the result. Under this criterion our method outperforms the strongest soft alternative on all three mazes ($2$--$5\times$), and the protocol identifies the remaining gap in that baseline's raw scores as routes through predicted transitions that do not exist. The same model executes in the real environment, placing second of seven against the published OGBench baselines on the hardest maze. Multimodal dynamics decide whether a JEPA world model can plan at all; a mixture of predictors with hard assignment is a minimal and verifiable fix.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!