SAM3D 기반 객체 중심 표현 정렬을 통한 비전-언어-액션 모델
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
비전-언어-액션 (VLA) 모델은 로봇 제어 분야에서 큰 잠재력을 보여주었지만, 대부분의 기존 모델은 2차원 시각-언어 기반 구조에 의존하며, 특히 가려짐, 자세 변화, 크기 변화 및 정밀한 공간 상호 작용 상황에서 대상 객체에 대한 세분화된 3차원 이해가 부족합니다. 본 연구에서는 $π_0$을 기반으로 구축된 객체 중심의 3차원 표현 정렬 프레임워크를 제안합니다. 구체적으로, SAM3D를 고정된 3차원 지도 모델로 사용하여 학습 과정에서 대상 객체의 3차원 정보를 제공합니다. 먼저, 객체 인식 모델을 사용하여 작업과 관련된 객체를 찾고, 해당 객체 마스크를 생성하고, SAM3D를 사용하여 객체 수준의 밀집된 3차원 표현을 추출합니다. 이 추출된 3차원 정보는 $π_0$의 중간 시각 특징과 정렬됩니다. 이를 통해 정책이 대상 객체의 3차원 정보를 학습하면서 원래의 RGB-언어-액션 추론 파이프라인을 유지할 수 있습니다. 또한, 테스트 단계에서 깊이 정보, 포인트 클라우드, 마스크, SAM3D 또는 추가적인 3차원 모듈이 필요하지 않습니다. 시뮬레이션 실험 결과, LIBERO 데이터셋에서 99.1%의 정확도를 달성했으며, CALVIN 데이터셋에서 평균 실행 길이가 4.11으로 향상되었습니다. 실제 환경에서의 실험을 통해 본 연구 방법론이 로봇이 여러 하위 작업에서 다양한 대상 객체에 집중해야 하는 장기적인 제어 시나리오에서 특히 효과적임을 확인했습니다.
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $π_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we localize task-relevant objects with object recognition models, generate corresponding object masks, and use SAM3D to extract dense object-level 3D representations, which are aligned with intermediate visual features of $π_0$. This enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time. Simulation experiments show consistent improvements, achieving 99.1\% on LIBERO and an average length of 4.11 on CALVIN. Real-world experiments further demonstrate that our method is particularly effective in long-horizon manipulation scenarios where the robot must focus on different target objects across multiple subtasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.