2606.30645v1 Jun 29, 2026 cs.RO

VLK: 재구성된 장면에서의 합성적 상호작용을 통해 학습하는 인간형 로봇의 위치 이동 및 조작

VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

K. Sreenath
K. Sreenath
Citations: 13,589
h-index: 55
Pieter Abbeel
Pieter Abbeel
Citations: 313
h-index: 9
Rocky Duan
Rocky Duan
Citations: 263
h-index: 7
Angjoo Kanazawa
Angjoo Kanazawa
Citations: 24,894
h-index: 60
Carmelo Sferrazza
Carmelo Sferrazza
Citations: 1,416
h-index: 21
Guanya Shi
Guanya Shi
Citations: 331
h-index: 9
Takara Truong
Takara Truong
Citations: 217
h-index: 3
Karen Liu
Karen Liu
Citations: 12
h-index: 2
Sirui Chen
Sirui Chen
Citations: 173
h-index: 5
Pei Xu
Pei Xu
Citations: 154
h-index: 4
Yen-Jen Wang
Yen-Jen Wang
Citations: 1,016
h-index: 11
Jiaman Li
Jiaman Li
Citations: 1,967
h-index: 17

인지 기반 인간형 로봇의 위치 이동 및 조작은 자율적인 관찰과 작업 지시를 전체 신체 움직임으로 연결해야 합니다. 이러한 매핑을 학습하려면 동기화된 자가 중심 이미지, 언어 명령, 그리고 로봇에 적합한 운동 궤적이 필요하지만, 현재까지 이러한 모든 요소를 대규모로 제공하는 데이터 소스는 존재하지 않습니다. 우리는 재구성된 장면에서 비전-언어-운동학(VLK)의 감독 신호를 합성적으로 생성함으로써 이 문제를 해결합니다. 우리의 파이프라인은 3D Gaussian Splatting을 사용하여 실제 크기의 실내 환경을 재구성하고, 우선적인 장면 정보를 활용하여 탐색 및 물체 상호 작용 운동 궤적을 합성하며, 이후에 쌍으로 연결된 자가 중심 관찰 데이터를 생성합니다. 우리는 인간의 개입 없이 48,000개의 쌍을 이루는 운동 궤적을 생성하고, 단기적인 전체 신체 운동 궤적을 예측하는 VLK 정책을 학습시켰습니다. 전체 신체 추적기는 이러한 예측을 실제 인간형 로봇의 행동으로 변환합니다. 우리는 물리적인 Unitree G1 로봇을 사용하여 탐색 및 단일 물체 운송 작업을 수행하며, 재구성된 장면에서의 합성적 상호작용이 시뮬레이션에서 실제 환경으로의 인지 기반 인간형 로봇 위치 이동 및 조작 학습에 효과적인 감독 신호를 제공한다는 것을 입증했습니다. 프로젝트 웹사이트: https://vision-language-kinematics.github.io/

Original Abstract

Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale. We address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. Our pipeline leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments, synthesizes navigation and object-interaction trajectories using privileged scene information, and renders paired egocentric observations after the fact. We produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts short-horizon whole-body kinematic trajectories. A whole-body tracker converts these predictions into actions on the physical humanoid. We evaluate on the physical Unitree G1 performing navigation and single-object transport, demonstrating that synthesized interactions in reconstructed scenes provide effective supervision for sim-to-real perception-based humanoid loco-manipulation. Project Website: https://vision-language-kinematics.github.io/

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!