2605.07278v1 May 08, 2026 cs.LG

예측은 가능하지만 계획은 어렵다: 잠재 세계 모델을 위한 RC-aux

Predictive but Not Plannable: RC-aux for Latent World Models

Guang Li
Guang Li
Citations: 490
h-index: 12
Keisuke Maeda
Keisuke Maeda
Citations: 622
h-index: 11
Takahiro Ogawa
Takahiro Ogawa
Citations: 2,559
h-index: 22
M. Haseyama
M. Haseyama
Citations: 3,565
h-index: 26
Wenyuan Li
Wenyuan Li
Citations: 32
h-index: 3

잠재 세계 모델은 정확한 단기 예측을 달성할 수 있지만, 동시에 계획에 적합하지 않은 잠재 공간을 유도할 수 있습니다. 주요 문제는 시공간적 불일치입니다. 이러한 모델은 종종 로컬 예측 감독 하에 훈련되지만, 유한한 행동 제약 조건 하에서 목표 지향적인 장기 검색을 위해 잠재 공간에서 사용됩니다. 이 불일치를 해결하기 위해, 본 논문에서는 재구성 없이 작동하는 잠재 세계 모델에 대한 경량 보정 방법인 Reachability-Correction (RC-aux)를 제안합니다. RC-aux는 세계 모델의 기본 구조를 변경하지 않고, 계획에 부합하는 감독 신호를 두 가지 축을 따라 추가합니다. 시간 축에서는 멀티 호라이즌 오픈 루프 예측을 통해 모델이 단일 단계 일관성뿐만 아니라 더 긴 시간 간격에 대한 예측을 학습하도록 합니다. 공간 축에서는 예산에 따른 도달 가능성 감독과 시간적 부정 샘플을 함께 사용하여, 잠재 공간이 궁극적으로 도달 가능한 상태와 현재 계획 호라이즌 내에서 도달 가능한 상태를 구별하도록 합니다. 테스트 시, 학습된 도달 가능성 신호는 또한 도달 가능성을 고려하는 계획기에 활용되어 목표 지향적이고 동시에 사용 가능한 예산 내에서 달성 가능한 경로를 선호하도록 할 수 있습니다. 본 논문에서는 LeWorldModel에 RC-aux를 적용하고, 지속적인 훈련 및 처음부터 훈련하는 두 가지 환경에서 성능을 평가했습니다. 목표 기반 픽셀 제어 작업 및 LIBERO-Goal 확장 작업에서, RC-aux는 LeWM 스타일의 계획 성능을 향상시키면서도 적은 추가 비용을 발생시킵니다. 이러한 결과는 잠재 세계 모델을 사용한 계획이 예측 정확도뿐만 아니라, 학습된 표현이 다운스트림 검색에 필요한 시간적 및 기하학적 구조를 인코딩하는지 여부에 달려 있음을 시사합니다. 코드 및 관련 정보는 https://github.com/Guang000/RC-aux 에서 확인할 수 있습니다.

Original Abstract

A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key issue is spatiotemporal mismatch: these models are often trained with local predictive supervision, but deployed for long-horizon goal-directed search in latent spaces where Euclidean distance may not reflect what is reachable within a finite action budget. We present the Reachability-Correction auxiliary objective (RC-aux), a lightweight correction for this mismatch in reconstruction-free latent world models. RC-aux keeps the world-model backbone unchanged and adds planning-aligned supervision along two axes. Along the time axis, multi-horizon open-loop prediction trains the model beyond one-step consistency. Along the space axis, budget-conditioned reachability supervision, together with temporal hard negatives, encourages the latent space to distinguish states that are eventually reachable from those reachable within the current planning horizon. At test time, the learned reachability signal can also be used by a reachability-aware planner to favor trajectories that are both goal-directed and attainable under the available budget. We instantiate RC-aux on LeWorldModel and evaluate it under both continuation-training and matched-from-scratch settings. Across goal-conditioned pixel-control tasks and a LIBERO-Goal extension, RC-aux improves LeWM-style planning with modest additional cost. These results suggest that planning with latent world models depends not only on predictive accuracy, but also on whether the learned representation encodes the temporal and geometric structure required by downstream search. The code is available at https://github.com/Guang000/RC-aux.

3 Citations
0 Influential
36.4657359028 Altmetric
13.9 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!