SG-WAM: 텍스트 기반 및 공간 인지 시맨틱 가이던스 - 월드-액션 모델
SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models
월드-액션 모델(WAM)은 로봇 조작 분야에서 유망한 패러다임으로 부상했습니다. 그러나 대부분의 기존 WAM은 언어적 지시 사항보다는 주로 시각적 단서에 의존하여 미래 영상과 행동을 생성합니다. 이는 오프더쉘 텍스트 인코더가 언어적 지시 사항을 시각적 관찰과 독립적으로 임베딩하기 때문입니다. 그 결과, 이러한 WAM이 예측하는 영상은 종종 해당 언어적 지시 사항과 의미적으로 일치하지 않아 예측된 행동의 정확도가 저하됩니다. 이러한 제한점을 극복하기 위해, 우리는 월드-액션 모델에 대한 시맨틱 가이드 방법인 SG-WAM을 제안합니다. SG-WAM은 비전-언어 모델(VLM)을 시맨틱 플래너로 활용하여 월드-액션 모델의 지시 사항 연관성을 강화합니다. 구체적으로, 우리는 VLM 기반 플래너를 훈련시켜 텍스트 기반 및 공간 인지 시맨틱 예측 기능을 제공하도록 합니다. 텍스트 기반 시맨틱 예측은 올바른 대상 객체를 식별하여 지시 사항을 연결하고, 공간 인지 시맨틱 예측은 정확한 조작을 위한 장면 기하 정보를 제공합니다. 그런 다음 이 예측 정보를 월드-액션 모델에 고수준 시맨틱 가이드로 주입하여 미래 영상 생성과 행동 예측이 언어적 지시 사항을 충실히 따르도록 합니다. 시뮬레이션 및 실제 환경에서의 광범위한 실험 결과, 우리의 시맨틱 가이드 방법이 우수한 성능을 보이며, 정밀한 조작 능력과 강력한 지시 사항 준수 능력을 보여줍니다.
World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual cues rather than language instructions, since off-the-shelf text encoders embed instructions independently of visual observations. As a result, the videos predicted by these WAMs are often semantically misaligned with their corresponding language instructions, which degrades the accuracy of the predicted actions. To overcome this limitation, we propose SG-WAM, a semantic guidance method for world-action models that leverages a vision-language model (VLM) as a semantic planner to enhance the instruction-grounding capacity of world-action models. Specifically, we train a VLM-based planner to predict text-grounded and spatial-aware semantic foresight. The text-grounded semantic foresight grounds the instruction by identifying the correct target objects, and the spatial-aware semantic foresight provides the scene geometry for precise manipulation. We then inject this foresight into the world-action model as high-level semantic guidance, ensuring that both future-video generation and action prediction faithfully follow the language instruction. Extensive experiments in simulation and the real world demonstrate the superiority of our semantic guidance method, showcasing precise manipulation and strong instruction-following capabilities.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.