2607.25912v1 Jul 28, 2026 cs.RO

SAM3D 기반 객체 중심 표현 정렬을 통한 비전-언어-액션 모델

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

Chinese Academy of Sciences
Chinese Academy of Sciences
Citations: 6,359
h-index: 35
Xiaoquan Sun
Xiaoquan Sun
Citations: 26
h-index: 3
Shan Jie
Shan Jie
Citations: 0
h-index: 0
Chen Cao
Chen Cao
Citations: 0
h-index: 0
Jiayu Kong
Jiayu Kong
Citations: 0
h-index: 0
Shenzhen Institute of Advanced Technology
Shenzhen Institute of Advanced Technology
Citations: 2
h-index: 1
Huazhong University of Science
Huazhong University of Science
Citations: 153
h-index: 7
Technology
Technology
Citations: 0
h-index: 0
Beijing University of Aeronautics
Beijing University of Aeronautics
Citations: 1
h-index: 1
Astronautics
Astronautics
Citations: 78
h-index: 4
Infiforce
Infiforce
Citations: 0
h-index: 0
Zonghe Liu
Zonghe Liu
Citations: 6
h-index: 1
Zetian Xu
Zetian Xu
Citations: 10
h-index: 2
Zongsheng Liu
Zongsheng Liu
Citations: 0
h-index: 0

비전-언어-액션 (VLA) 모델은 로봇 제어 분야에서 큰 잠재력을 보여주었지만, 대부분의 기존 모델은 2차원 시각-언어 기반 구조에 의존하며, 특히 가려짐, 자세 변화, 크기 변화 및 정밀한 공간 상호 작용 상황에서 대상 객체에 대한 세분화된 3차원 이해가 부족합니다. 본 연구에서는 $π_0$을 기반으로 구축된 객체 중심의 3차원 표현 정렬 프레임워크를 제안합니다. 구체적으로, SAM3D를 고정된 3차원 지도 모델로 사용하여 학습 과정에서 대상 객체의 3차원 정보를 제공합니다. 먼저, 객체 인식 모델을 사용하여 작업과 관련된 객체를 찾고, 해당 객체 마스크를 생성하고, SAM3D를 사용하여 객체 수준의 밀집된 3차원 표현을 추출합니다. 이 추출된 3차원 정보는 $π_0$의 중간 시각 특징과 정렬됩니다. 이를 통해 정책이 대상 객체의 3차원 정보를 학습하면서 원래의 RGB-언어-액션 추론 파이프라인을 유지할 수 있습니다. 또한, 테스트 단계에서 깊이 정보, 포인트 클라우드, 마스크, SAM3D 또는 추가적인 3차원 모듈이 필요하지 않습니다. 시뮬레이션 실험 결과, LIBERO 데이터셋에서 99.1%의 정확도를 달성했으며, CALVIN 데이터셋에서 평균 실행 길이가 4.11으로 향상되었습니다. 실제 환경에서의 실험을 통해 본 연구 방법론이 로봇이 여러 하위 작업에서 다양한 대상 객체에 집중해야 하는 장기적인 제어 시나리오에서 특히 효과적임을 확인했습니다.

Original Abstract

Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $π_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we localize task-relevant objects with object recognition models, generate corresponding object masks, and use SAM3D to extract dense object-level 3D representations, which are aligned with intermediate visual features of $π_0$. This enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time. Simulation experiments show consistent improvements, achieving 99.1\% on LIBERO and an average length of 4.11 on CALVIN. Real-world experiments further demonstrate that our method is particularly effective in long-horizon manipulation scenarios where the robot must focus on different target objects across multiple subtasks.

0 Citations
0 Influential
17.5 Altmetric
87.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!