SUNTA: 놀람 기반 청킹을 활용한 계층적 비디오 예측
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking
계층적 상태 공간 모델(HSSM)은 시퀀스를 시간 단위로 분할하여 장기 예측에 유망한 접근 방식을 제공합니다. 그러나 이러한 모델의 성능은 청크 경계를 어떻게 결정하는지에 크게 의존합니다. 기존의 HSSM은 일반적으로 고정 길이 청킹 또는 유사성 기반 경계 탐지를 사용하지만, 이러한 방법은 종종 데이터의 내재적인 시간 구조와 일치하지 않습니다. 우리는 청킹이 예측 오류에 의해 주도되어야 한다고 주장하는데, 이는 장거리 컨텍스트가 필요할 때를 보다 직접적으로 나타냅니다. 그러나 놀람 기반 청킹을 HSSM에 통합하는 것은 전체 훈련 과정에서의 계층적 붕괴 및 개방 루프 예측 시 놀람 신호의 부재와 같은 중요한 과제를 야기합니다. 이러한 문제를 해결하기 위해, 우리는 놀람 기반 중첩 시간 추상화(SUNTA)라는 방법을 제안합니다. SUNTA는 분리된 학습 전략을 사용하여 놀람 신호를 유지하고, 상상 속의 실행 과정을 통해 내부 불일치를 최상위 레벨의 놀람 지표로 활용하여 청크 경계를 결정합니다. 2D 및 3D 환경에서의 비디오 예측 작업에 대한 실험 결과, SUNTA는 기존 모델보다 뛰어난 성능을 보이며, 250개의 타임스텝 동안 정확한 예측을 유지하는 반면, 모든 기존 모델은 처음 10개 타임스텝 내에서 성능이 저하됩니다.
Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks. However, their performance hinges on how chunk boundaries are determined. While prior HSSMs typically rely on fixed-length chunking or similarity-based boundary detection, these methods often misalign with the intrinsic temporal structure of the data. We argue that chunking should instead be driven by prediction errors, which more directly indicate when longer-range context becomes necessary. Nevertheless, integrating surprise-based chunking into HSSMs introduces critical challenges, including hierarchical collapse during end-to-end training and the absence of surprise signals during open-loop prediction. To address these issues, we propose Surprise-based Nested Temporal Abstraction (SUNTA), a method that employs a decoupled training strategy to preserve surprise signals and uses internal inconsistency as a top-down surprise metric to determine chunk boundaries within imagined rollouts. Experiments on video prediction tasks in 2D and 3D environments demonstrate that SUNTA outperforms baselines, uniquely maintaining accurate predictions over 250 timesteps, whereas all baselines degrade within the first 10 timesteps.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.