ST-WAM: 시각적 분포 변화에 강건한 조작을 위한 의미론적-시간적 세계 동작 모델
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
세계 동작 모델(WAM)은 로봇의 행동과 미래의 시각적 동역학을 동시에 모델링하여 유망한 패러다임으로 떠오르고 있습니다. 그러나 기존 WAM은 픽셀 기반 미래 예측에 의존하기 때문에, 이는 행동과 관련된 상태 변화를 작업과 관련 없는 시각 정보와 혼동시켜 시각적 분포 변화에 대한 강건성을 제한할 수 있습니다. 본 연구에서는 '훈련 데이터 환각(Training-Distribution Hallucination)'이라는 현상을 규명했습니다. 이 현상은 시각적으로 변화된 관찰을 기반으로 미래를 예측할 때, 실제 장면과 일치하는 내용 대신 훈련 데이터의 내용을 반영하는 경향이 있습니다. 제어된 프레임 삼중항 분석 결과, DINOv3 특징은 시각적 변화에 더 안정적이며, Wan-VAE 잠재 벡터보다 작업 상태 구분을 더 잘 유지하는 것으로 나타났습니다. 예측된 미래를 수정하는 대신, 본 연구에서는 의미론적-시간적 WAM(ST-WAM)을 제안합니다. ST-WAM은 DINOv3를 사용하여 미래 예측 및 과거 정보 검색을 위한 공유된 의미론적 표현을 제공하고, 동시에 세밀한 VAE 동역학을 유지합니다. 이 모델의 '이중 공간 미래 전문가(Dual-Space Future Experts)'는 미래의 VAE 잠재 벡터와 DINO 특징을 동시에 예측하며, '현재 기준 의도 검색(Current-Anchored Intent Retrieval)'은 현재의 시각-언어 맥락 하에서 작업과 관련된 정보를 최근의 DINO 과거 정보에서 검색합니다. ST-WAM은 추가적인 몸체 사전 훈련 또는 작업별 주석 없이 end-to-end 방식으로 학습되며, 추론 과정에서 명시적인 미래 생성이 필요하지 않습니다. LIBERO 데이터셋에서 98.7%, RoboTwin 2.0 데이터셋에서 92.8%의 성능을 달성했으며, 특히 Fast-WAM과 비교하여 zero-shot LIBERO-Plus 성능이 21.3%p 향상되었고, 시각적 변화가 있는 환경에서의 실제 성공률이 25.8%에서 61.5%로 두 배 이상 증가했습니다. 이러한 결과는 의미론적-시간적 모델링이 픽셀 기반 동역학을 효과적으로 보완하여 강건한 조작을 가능하게 한다는 것을 보여줍니다.
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.