FF-JEPA: 잠재 계획기를 활용한 월드 모델에서의 장기 계획
FF-JEPA: Long-Horizon Planning in World Models with Latent Planners
합쳐진 임베딩 예측 아키텍처(JEPAs)는 유망한 월드 모델링 능력을 보여주며, Cross-Entropy Method (CEM)와 같은 방법을 사용하여 액션 경로를 최적화함으로써 잠재 공간에서 계획을 가능하게 합니다. 그러나 이러한 방법은 계산 비용이 너무 많이 들고 장기 계획에 효과적이지 않습니다. 또한, 이러한 방법들은 일반적으로 목표 상태의 명시적인 이미지를 필요로 하지만, 이는 실제 작업에서는 항상 가능한 것은 아닙니다. 본 연구에서는 이러한 한계를 해결하기 위해 Forward-Forward-JEPA (FF-JEPA)라는 계층적 접근 방식을 제안합니다. 표준 액션 기반 예측 모델 외에도, 현재 상태를 기반으로 다음 하위 목표를 예측하는 액션이 없는 잠재 계획기를 도입했습니다. 이 방법은 목표 이미지의 필요성을 없애고 복잡한 경로를 일련의 관리 가능한 단기 최적화 문제로 분해하여 장기 계획을 가능하게 합니다. PushT 환경에서의 예비 결과는 FF-JEPA가 기존 월드 모델의 장기적인 성능 저하 문제를 성공적으로 극복한다는 것을 보여주며, 이는 목표 정보 없이 계획하는 데 있어 유망한 방향임을 시사합니다.
Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods like the Cross-Entropy Method (CEM). These methods are, however, too computationally expensive and ineffective for long-horizon planning. Furthermore, these methods typically require an explicit image of the goal state, which is not always possible in real-world tasks. In this work, we tackle these limitations by proposing Forward-Forward-JEPA (FF-JEPA), a hierarchical approach leveraging two forward dynamics models. Alongside a standard action-conditioned forward model, we introduce an action-free latent planner that predicts the next subgoal given the current state. This approach removes the need for goal images and enables long-horizon planning by decomposing complex trajectories into a sequence of tractable, short-term optimization problems. Preliminary results on PushT demonstrate that FF-JEPA successfully overcomes flat world models' long-horizon collapse, highlighting this approach as a promising direction for goal-free planning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.