2608.05042v1 Aug 05, 2026 cs.RO

BridgeVLA++: 데이터 효율적이고 일반화 가능하며 메모리 기반의 3차원 로봇 조작을 위한 비전-언어-액션 프레임워크

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

Tieniu Tan
Tieniu Tan
Citations: 258
h-index: 7
Liang Wang
Liang Wang
Citations: 133
h-index: 4
Peiyan Li
Peiyan Li
Citations: 264
h-index: 7
Xiao Ma
Xiao Ma
Citations: 861
h-index: 13
Yuan Xu
Yuan Xu
Citations: 47
h-index: 4
Yan Huang
Yan Huang
Citations: 39
h-index: 4
Yixiang Chen
Yixiang Chen
Citations: 96
h-index: 4
Qisen Ma
Qisen Ma
Citations: 41
h-index: 2
He Guan
He Guan
Citations: 460
h-index: 7
Tao Kong
Tao Kong
Citations: 43
h-index: 2
Jiabing Yang
Jiabing Yang
Citations: 33
h-index: 3
Yuze Zhu
Yuze Zhu
Citations: 0
h-index: 0
Hongtao Wu
Hongtao Wu
Citations: 1,339
h-index: 9

사전 학습된 비전-언어 모델(VLMs)을 활용하여 비전-언어-액션(VLA) 모델을 구축하는 것은 3차원 로봇 조작 분야에서 유망한 접근 방식입니다. 그러나 기존의 3차원 VLA 방법은 여전히 많은 데이터를 필요로 하며, 데이터 분포 변화에 대한 일반화 능력이 제한적이고, 과거 관찰 기록을 명시적으로 기억하지 못합니다. 이러한 한계는 데이터가 부족하거나, 개방형 환경이며, 과거 정보 의존성이 높은 조작 시나리오에서의 활용을 어렵게 만듭니다. 저희의 이전 연구인 BridgeVLA는 3차원 액션 학습 과정에서 사전 학습된 VLM의 입력-출력 정렬을 유지함으로써 데이터 효율성과 일반화 능력을 향상시켰습니다. 구체적으로, 원본 점군 데이터를 멀티뷰 이미지로 투영하고 로봇 동작을 생성하기 전에 중간 히트맵을 예측합니다. 이번 연구에서는 BridgeVLA에 통합된 시공간 메모리 아키텍처를 추가하여 BridgeVLA++를 개발했습니다. 이 새로운 아키텍처는 지속적인 공간적 맥락과 시간적 상호 작용 기록을 모델링하며, BridgeVLA의 데이터 효율성과 일반화 능력을 유지하면서 과거 관찰 기록에 대한 추론이 가능합니다. 광범위한 실험 결과, 저희 프레임워크는 공간 조작 작업에서 뛰어난 성능을 보이며 강력한 일반화 능력을 보여줍니다. 또한, BridgeVLA++는 원본 BridgeVLA의 데이터 효율성과 일반화 능력 저하 없이 두 가지 어려운 메모리 의존적인 조작 벤치마크에서 최고 수준의 성능을 달성했습니다. 더불어, BridgeVLA++는 양손 조작 환경에서도 효과적으로 작동하며, 추가적인 실제 로봇 플랫폼에서도 검증되어 다양한 작업, 환경 및 로봇 플랫폼에 대한 확장성을 입증합니다. 이러한 결과들은 BridgeVLA++가 데이터 효율적인 학습, 강력한 일반화, 그리고 효과적인 메모리 기반 로봇 조작을 동시에 지원하는 통합된 3차원 비전-언어-액션 프레임워크임을 보여줍니다. 프로젝트 웹사이트: https://bridgevla-plus.github.io/.

Original Abstract

Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!