G$^3$VLA: 시각-언어-행동 모델을 위한 기하학적 유도 편향
G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
시각-언어-행동 (VLA) 모델은 사전 학습된 시각-언어 기반 모델의 의미론적 지식을 활용하여 일반적인 로봇 조작 분야에서 빠른 발전을 이루었지만, 이러한 모델의 시각 토큰은 여전히 로봇 카메라의 보정된 기하학적 정보가 아닌 2D 이미지 좌표에 기반합니다. 이는 특히 다중 카메라 환경에서 더욱 두드러지는데, 각 뷰는 알려진 내부 및 외부 파라미터로 연결되어 있지만 독립적인 이미지로 처리됩니다. 본 연구에서는 사전 학습된 VLA 모델의 시각 토큰 스트림에 보정된 구조를 주입하는 카메라 인식 기하학적 모듈인 G$^3$VLA를 제안합니다. 이는 액션 공간이나 모방 목표를 변경하지 않고, 내부 조건이 적용된 레이 임베딩, 투영 위치 인코딩 (PRoPE), 그리고 양방향 교차 뷰 융합을 결합합니다. 기하학적 지도 학습은 필요에 따라 실제 포인트 맵 또는 신뢰도 기반 $π^3$X 교사 모델의 예측을 통해 제공되며, 별도의 깊이 센서나 수동 주석 없이 구현됩니다. $π_0$ 데이터셋에서 G$^3$VLA를 적용한 결과, LIBERO 스위트, RoboCasa24, RoboTwin2.0 및 실제 로봇 환경에서 일관된 성능 향상을 보였으며, 특히 공간적이고 객체에 민감한 작업에서 가장 큰 개선 효과를 나타냈습니다. 또한 $π_{0.5}$ 및 GR00T 1.5 데이터셋에서도 검증을 수행했는데, 그 결과는 기하학적으로 인식하는 토큰이 직접 액션 생성 경로에 접근할 때 기하학적 전이가 가장 효과적임을 시사합니다. 프로젝트 웹 페이지는 https://sites.google.com/view/g3vla 입니다.
Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their visual tokens remain grounded in 2D image coordinates rather than the calibrated geometry of the robot's cameras -- a mismatch especially pronounced in multi-camera setups, where views are coupled by known intrinsics and extrinsics yet processed as independent images. We propose G$^3$VLA, a camera-aware geometric module that injects calibrated structure into the visual-token stream of a pretrained VLA without altering its action space or imitation objective, combining intrinsic-conditioned ray embeddings, projective positional encoding (PRoPE), and bidirectional cross-view fusion. Geometric supervision is provided either from ground-truth point maps when available, or from confidence-gated $π^3$X teacher predictions, requiring no depth sensors or manual annotations. Instantiated on $π_0$, G$^3$VLA yields consistent gains across the LIBERO suites, RoboCasa24, RoboTwin2.0, and real-robot settings, with the largest improvements on spatially and object-sensitive tasks. We further validate on $π_{0.5}$ and GR00T 1.5, with results suggesting that geometric transfer is most effective when geometry-aware tokens have direct access to the action generation pathway. Our project page is at https://sites.google.com/view/g3vla
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.