2606.31329v1 Jun 30, 2026 cs.RO

3D HAMSTER: 3차원 경로 지침을 통한 계층적 시각-언어-행동 모델에서 계획과 제어를 연결

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

Hojoon Lee
Hojoon Lee
Citations: 311
h-index: 10
Jaegul Choo
Jaegul Choo
Citations: 366
h-index: 11
Dongyoon Hwang
Dongyoon Hwang
Citations: 170
h-index: 8
Hyojin Jang
Hyojin Jang
Citations: 33
h-index: 3
Hoiyeong Jin
Hoiyeong Jin
Citations: 22
h-index: 3
Jueun Mun
Jueun Mun
Citations: 6
h-index: 1
Minho Park
Minho Park
Citations: 13
h-index: 2
Hyunseung Kim
Hyunseung Kim
Citations: 429
h-index: 8
Byungkun Lee
Byungkun Lee
Citations: 20
h-index: 2
Dongjin Kim
Dongjin Kim
Citations: 4
h-index: 1

계층적 시각-언어-행동(VLA) 모델은 로봇 조작의 일반화 성능 향상을 위해 고수준 계획과 저수준 제어를 분리합니다. 최근 연구에서는 시각-언어 모델(VLM)이 예측하는 2차원 엔드 이펙터 경로를 하위 정책에 대한 명시적인 지침으로 사용합니다. 그러나 최첨단 저수준 정책은 포인트 클라우드를 기반으로 3차원 메트릭 공간에서 작동하며, 깊이 정보가 부족한 2차원 지침을 제공하면 각 위치에 해당 장면 표면 아래의 깊이가 할당되어 기하학적으로 왜곡된 경로가 생성됩니다. 본 연구에서는 계획기가 직접 측정 가능한 신뢰성 있는 3차원 경로를 출력하도록 하는 계층적 프레임워크인 3D HAMSTER를 제안합니다. VLM에 전용 깊이 인코더와 밀집 깊이 재구성 목표를 추가하여 3차원 위치 시퀀스를 예측하고, 이를 포인트 클라우드 기반의 저수준 정책에 직접 통합합니다. 3D HAMSTER는 3차원 경로 예측, 시뮬레이션 및 실제 로봇 조작 환경에서 기존 VLM과 2차원 지침 기반 모델보다 우수한 성능을 보이며, 특히 외관 변화, 새로운 언어, 공간적 조건 및 시각적 조건 하에서 더 큰 성능 향상을 나타냅니다. 프로젝트 페이지는 https://davian-robotics.github.io/3D_HAMSTER/ 에서 확인할 수 있습니다.

Original Abstract

Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a pointcloudbased low-level policy. Across 3D trajectory prediction, simulation, and real-world manipulation, 3D HAMSTER consistently outperforms proprietary VLMs and 2D-guided baselines, with the largest gains under appearance-altering shifts and unseen language, spatial, and visual conditions. The project page is available at https://davian-robotics.github.io/3D_HAMSTER/.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!