2608.01826v1 Aug 03, 2026 cs.RO

멀티뷰 통합 카메라 필드: RGB 데이터만 사용한 다중 카메라 VLA 정책을 위한 기하학적 구조 기반 액션 표현

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies

Yehao Lu
Yehao Lu
Citations: 195
h-index: 5
Pei Lin
Pei Lin
Citations: 24
h-index: 3
Jiarui Yang
Jiarui Yang
Citations: 0
h-index: 0
Yuning Su
Yuning Su
Citations: 0
h-index: 0
Yufeng Xie
Yufeng Xie
Citations: 0
h-index: 0
Yu Zhong
Yu Zhong
Citations: 24
h-index: 3
Haiyu Lan
Haiyu Lan
Citations: 0
h-index: 0
Tianjing Hao
Tianjing Hao
Citations: 0
h-index: 0
Kaixiang Lu
Kaixiang Lu
Citations: 0
h-index: 0
Chuang Wang
Chuang Wang
Citations: 0
h-index: 0
Enyu Li
Enyu Li
Citations: 0
h-index: 0
Junwei Liang
Junwei Liang
Citations: 273
h-index: 10

비전-언어-액션(VLA) 모델은 로봇 조작 분야에서 강력한 일반화 능력을 보여주지만, 복잡하고 접촉이 많은 작업은 종종 엔드 이펙터, 물체 및 가려진 목표를 동시에 포착하는 다중 카메라 관찰로부터 더 큰 이점을 얻습니다. 기존의 다중 카메라 VLA 모델은 일반적으로 뷰 토큰을 연결하지만, 이는 액션 표현의 메트릭 깊이를 약화시키고 카메라 간에 일관성을 유지하기 어렵게 만듭니다. 본 연구에서는 Multi-View Unified Camera Fields (MVUCF)라는 학습 전용 프레임워크를 제안합니다. MVUCF는 뷰 간에 공유되는 액션 지향 잠재 필드를 형성합니다. 좌표 기반 깊이 추정 목표는 메트릭 깊이를 복구하는 데 도움이 되며, 사전 처리 인지 대응성 목표는 서로 다른 카메라에서 동일한 물리적 지점을 관찰하는 토큰을 정렬합니다. 이 두 가지 요소 모두 액션 모듈에 의해 소비되는 은닉 상태를 직접적으로 형성합니다. 기하학 정보를 주입한 후에는 깊이 정보, 카메라 보정 및 추가적인 헤드를 제거하여, 배포 시 원래의 RGB 데이터만 사용하는 그래프 구조에서 추가적인 추론 연산(FLOPs) 없이 사용할 수 있습니다. 독립적인 검증 결과, MVUCF는 더 강력한 깊이 복구 능력과 뷰 간 일치성을 보여줍니다. GR00T-N1.6 환경에서, MVUCF는 LIBERO 데이터셋에서 98.9%의 성능을 달성하고, LIBERO-Plus 데이터셋에서 22.4점의 성능 향상을 보이며, 터치, 이동 및 배치, 접촉 상호작용 등 세 가지 액션 유형에 걸쳐 여섯 가지 RoboTwin 작업에서 성공률을 23.3점 향상시켰습니다. 실제 휴머노이드 로봇 실험 결과는 RGB 데이터만 사용하는 환경에서도 MVUCF가 실질적인 효과를 가진다는 것을 입증합니다.

Original Abstract

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-camera observations that jointly capture the end effector, objects, and targets under occlusion. Existing multi-camera VLAs usually concatenate view tokens, leaving action representations weak in metric depth and inconsistent across cameras. We introduce Multi-View Unified Camera Fields (MVUCF), a training-only framework that forms a shared action-facing latent field across views. A coordinate-query depth objective makes metric depth recoverable, while a preprocessing-aware correspondence objective aligns tokens observing the same physical point from different cameras. Both directly shape the hidden states consumed by the action module. After geometry injection, depth, camera calibration, and auxiliary heads are removed, so deployment uses the original RGB-only graph with no extra inference FLOPs. Held-out probes confirm stronger depth recovery and cross-view matching. Under matched GR00T-N1.6 settings, MVUCF reaches 98.9% on LIBERO, improves LIBERO-Plus by 22.4 points, and raises success by 23.3 points across six RoboTwin tasks spanning three action families: touch, move-and-place, and contact interaction. Real-world humanoid experiments further provide evidence of its practical effectiveness under RGB-only deployment.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!