SAGE: 잠재 세계 모델 기반 계획을 위한 하위 목표 조건부 행동 생성
SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
잠재 세계 모델은 행동에 조건을 부여한 예측 동역학을 학습하고 이를 내부 시뮬레이터로 사용하여 후보 행동 시퀀스를 상상하고 평가하는 강력한 계획 패러다임으로 부상했습니다. 그러나 계획 범위가 증가함에 따라 성능은 제안의 품질에 의해 점점 더 제한됩니다. 고정된 후보 예산으로 지수적으로 더 큰 행동 공간을 탐색해야 하므로, 세계 모델이 평가를 위해 양질의 후보 미래를 학습하기 어렵습니다. 본 논문에서는 무작위 제안 초기화를 대체하는 사전 조건부 계획자를 소개합니다. 각 계획 단계에서 목표에 조건을 부여한 생성기는 특정 기간 동안 도달 가능한 다음 잠재 하위 목표를 예측하며, 이는 후보 행동 시퀀스 생성을 조건화하는 데 사용됩니다. 시간 척도 전체의 의미 정보를 포착하기 위해 다양한 기간의 하위 목표를 사전 정보로 사용하여 미세한 로컬 제어와 고차원 장기 진행 간의 균형을 맞춥니다. 그런 다음 동결된 세계 모델은 실행 전에 이러한 하위 목표 조건부 제안을 평가하고 개선합니다. PushT 및 OGBench Cube에 대한 실험 결과, 잠재 하위 목표 분해를 사전 조건부 행동 생성과 결합하면 장기 계획 성능이 크게 향상되는 동시에 우수한 단기 성능을 유지할 수 있음을 보여줍니다. 구체적으로, 대상 오프셋이 $150$인 경우 PushT의 성공률이 $12.7%$에서 $64.7%$로, OGBench Cube의 성공률이 $26.7%$에서 $67.3%$로 향상되었습니다.
Latent world models have emerged as a powerful planning paradigm by learning action-conditioned predictive dynamics and using them as internal simulators to imagine and evaluate candidate action sequences. However, as the planning horizon grows, performance becomes increasingly constrained by proposal quality: a fixed candidate budget must search an exponentially larger action space, making it difficult to expose the world model to high-quality candidate futures for evaluation. In this paper, we introduce a prior-conditioned planner that replaces random proposal initialization with structured guidance. At each planning stage, a goal-conditioned generator predicts the next reachable latent subgoal for a specified duration, which is then used to condition the generation of candidate action sequences. To capture semantic information across temporal scales, we use subgoals of varying durations as priors, balancing fine-grained local control with higher-level long-horizon progress. Then the frozen world model evaluates and refines these subgoal-conditioned proposals before execution. Experiments on PushT and OGBench Cube show that coupling latent subgoal decomposition with prior-conditioned action generation substantially improves long-horizon planning while preserving strong short-horizon performance. To be specific, when the target offset is $150$, it raises PushT success from $12.7\%$ to $64.7\%$ and OGBench Cube success from $26.7\%$ to $67.3\%$.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.