2608.05903v1 Aug 06, 2026 cs.CV

Robust-WAM: 월드-액션 모델에서 생성적 사전 훈련과 의미 기반 예측 능력 간의 연결

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Junjie He
Junjie He
Citations: 81
h-index: 3
Wenxuan Song
Wenxuan Song
Citations: 697
h-index: 14
Zhide Zhong
Zhide Zhong
Citations: 165
h-index: 6
Haodong Yan
Haodong Yan
Citations: 123
h-index: 5
Haoang Li
Haoang Li
Citations: 401
h-index: 9
MingMing Yu
MingMing Yu
Citations: 0
h-index: 0
Jiadi You
Jiadi You
Citations: 0
h-index: 0
Yingjie Cai
Yingjie Cai
Citations: 96
h-index: 5
Xu Yan
Xu Yan
Citations: 165
h-index: 8
Jiaguang Zhu
Jiaguang Zhu
Citations: 44
h-index: 3
Yangyang Zheng
Yangyang Zheng
Citations: 42
h-index: 1
Yuqiao Du
Yuqiao Du
Citations: 0
h-index: 0
Guanyi Zhao
Guanyi Zhao
Citations: 28
h-index: 4
Bingbing Liu
Bingbing Liu
Citations: 193
h-index: 4

주류 월드-액션 모델(WAM)은 로봇 제어를 위해 사전 훈련된 비디오 생성 모델(VGM)을 활용하며, 학습된 역학적 지식을 액션 예측에 적용합니다. 이러한 VGM은 일반적으로 변분 오토인코더(VAE)의 잠재 공간에서 훈련됩니다. 그러나 VAE 잠재 공간은 주로 픽셀 재구성을 최적화하는데, 이는 미세한 외형 디테일에 대한 높은 성능을 제공하지만 시각적인 변화에 취약합니다. 최근 연구에서는 외형 변화에 더 강건한 의미 기반 잠재 공간을 활용하는 WAM을 개발하고 있습니다. 하지만 이러한 모델은 VAE 공간에서만 가능한 대규모 VGM 사전 훈련의 장점을 활용할 수 없습니다. 이러한 어려움을 해결하기 위해, 우리는 Robust-WAM이라는 일반적인 후처리 방법을 제안합니다. 이 방법은 비디오 생성 기반 WAM에 적용되며, VAE 기반의 생성 경로를 유지하면서 액션 스트림에 경량화된 의미 기반 예측 정렬 목표를 추가합니다. 이를 통해 대규모 VGM 사전 훈련의 장점을 유지하면서 외형 변화에 영향을 받지 않는 역학적 특성을 확보하여 조명 변화 및 기타 시각적인 이상 상황에서도 안정적인 성능을 보장합니다. 구체적으로, 우리는 학습 가능한 쿼리 토큰을 사용하여 미래 장면의 의미 정보를 액션 스트림으로 가져오고, 이들의 출력 은닉 상태를 미래 프레임의 의미 기반 예측 정보와 정렬합니다. 각 쿼리와 해당 쿼리가 설명하는 미래 시점 사이의 시간적 대응 관계를 설정하기 위해, 일치하는 액션 토큰의 위치 인코딩을 사용합니다. 이상 검출 및 일반화 시뮬레이션 벤치마크, 그리고 실제 로봇 환경에서의 실험 결과는 Robust-WAM이 기존 WAM 모델들의 성공률을 지속적으로 향상시키면서도 정상적인 데이터셋에서의 성능 저하를 일으키지 않는다는 것을 보여줍니다.

Original Abstract

Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs are typically trained in a variational autoencoder (VAE) latent space. However, the VAE latent space is optimized for pixel reconstruction, which rewards fine appearance detail and leaves the action prediction fragile under visual shifts. Recent works build WAMs in semantic latent space, which are more robust to appearance shifts. However, these models cannot leverage the large-scale VGM pretraining that exists only in VAE space. To overcome this dilemma, we propose Robust-WAM, a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream. This retains the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics that stay reliable under illumination shifts and other visual out-of-distribution conditions. Specifically, we employ learnable query tokens to bring future-scene semantics into the action stream by aligning their output hidden states with the semantic foresight of future ground-truth frames. To establish the temporal correspondence between each query and the future step it describes, we give it the positional encoding of the matching action tokens. Experiments on out-of-distribution generalization simulation benchmarks and a real-robot setup show that our Robust-WAM consistently improves the success rates of multiple WAM baselines without sacrificing in-distribution performance.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!