2608.06729v1 Aug 07, 2026 cs.RO

AtlasVLA: 비전-언어-행동 모델을 위한 지속적인 세계-자기 상태 모델링

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

Longteng Guo
Longteng Guo
Citations: 1,611
h-index: 18
Xingjian He
Xingjian He
Citations: 615
h-index: 11
Zilin Zhu
Zilin Zhu
Citations: 380
h-index: 3
Yanghong Mei
Yanghong Mei
Citations: 6
h-index: 1
Jing Liu
Jing Liu
Citations: 56
h-index: 5
Guiyu Zhao
Guiyu Zhao
Citations: 113
h-index: 5
Yu Zhang
Yu Zhang
Citations: 26
h-index: 2
Bin Cao
Bin Cao
Citations: 49
h-index: 4
Mingming Yu
Mingming Yu
Citations: 37
h-index: 4
Jie Jiang
Jie Jiang
Citations: 190
h-index: 5

비전-언어-행동(VLA) 모델은 로봇 공학 분야의 발전에 기여했지만, 근본적으로 반응적인 패러다임 때문에 부분 관찰 환경과 장기적인 작업에서 성능이 제한됩니다. 특히 손목에 부착된 단일 카메라로만 작동할 때, 물체가 시야에서 벗어지면서 발생하는 인지 능력 저하 및 다단계 실행 과정에서의 작업 진행 상황 망각 문제가 발생합니다. 이러한 문제점을 해결하기 위해, 우리는 지속적인 세계-자기 상태를 기반으로 직접적인 반응형 조작에서 적극적인 추론으로 전환하는 새로운 프레임워크인 AtlasVLA를 제안합니다. AtlasVLA는 4차원 영구적 세계 상태 메모리와 자기-작업 상태 메모리라는 이중 메모리 구조를 특징으로 합니다. 영구적 세계 상태 메모리는 일시적인 2차원 관찰 데이터를 전역적으로 업데이트된 복셀 해시 기반 공간 상태로 변환하여 시각적 사각지대를 해결하며, 자기-작업 상태 메모리는 과거의 자기 상태와 작업 진행 상황을 추적합니다. 이러한 세계-자기 상태 정보를 활용하여 설계된 디퓨전 트랜스포머(DiT)는 강력한 공간 추론 능력을 제공합니다. LIBERO, RLBench 및 실제 환경 벤치마크를 통한 광범위한 실험 결과, AtlasVLA는 단일 손목 카메라만 사용하여 최첨단 성능을 달성함을 보여줍니다. 특히, 다양한 관점을 사용하는 기존 모델보다 훨씬 뛰어난 성능을 보이며, LIBERO-Long 데이터셋에서 절대 성공률이 9.4% 향상되고, 실제 환경의 장기적인 작업에서 17.5% 개선되었습니다.

Original Abstract

While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!