2606.10267v1 Jun 09, 2026 cs.RO

로봇 정책 조율에서 중요한 요소: 계층적 시각-언어-행동 에이전트에 대한 체계적인 연구

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents

Jiaheng Hu
Jiaheng Hu
Citations: 1,777
h-index: 13
Mohit Shridhar
Mohit Shridhar
Citations: 4,510
h-index: 16
Caden Lu
Caden Lu
Citations: 3,294
h-index: 2
Dhruv Shah
Dhruv Shah
Citations: 8,455
h-index: 29
Hao-Tien Lewis Chiang
Hao-Tien Lewis Chiang
Citations: 544
h-index: 5
Jie Tan
Jie Tan
Citations: 1,550
h-index: 7
Annie Xie
Annie Xie
Citations: 3,684
h-index: 5

계층적 시각-언어-행동(Hi-VLA) 시스템은 고수준 VLM 플래너를 사용하여 작업을 언어 기반의 하위 목표로 분해하고, 저수준 VLA 제어기가 이를 실행함으로써 복잡한 로봇 조작에 대한 유망한 패러다임으로 부상했습니다. 최근의 실증적인 발전에도 불구하고, 이러한 시스템을 위한 통일된 설계 원칙은 부족합니다. 기존의 Hi-VLA 시스템들은 플래너와 제어기의 선택 및 연결 방식, 두 모드 간 전환 메커니즘, 그리고 플래너 내에서의 관찰 및 메모리 표현 방식에서 차이를 보입니다. 본 논문에서는 로봇 조작을 위한 Hi-VLA 설계에 대한 체계적인 연구를 제시합니다. 대표적인 Hi-VLA 에이전트를 옵션 기반 제어 프레임워크 하에 통합하고, 단기, 장기 및 추론 집약적인 작업에서 핵심 설계 선택 사항을 비교 분석했습니다. 우리의 분석은 Hi-VLA 시스템 구축을 위한 실질적인 원칙을 도출하며, 모델 선택과 인터페이스 메커니즘이 성능에 공동으로 미치는 영향을 보여줍니다. 이러한 원칙을 적용하면 시뮬레이션 환경 및 실제 ALOHA 로봇에서의 실험에서 평면 VLA 제어 또는 단순하게 설계된 계층 구조보다 훨씬 강력한 시스템을 구축할 수 있습니다. 종합적으로, 우리의 결과는 더욱 능숙하고 견고하며 체계적인 계층적 VLA 에이전트 개발을 위한 기반을 제공합니다. 자세한 정보 및 영상은 jiahenghu.github.io/hi-vla 에서 확인할 수 있습니다.

Original Abstract

Hierarchical vision-language-action (Hi-VLA) systems have emerged as a promising paradigm for complex robot manipulation, by using high-level VLM planners to decompose tasks into language subgoals executed by low-level VLA controllers. Despite recent empirical progress, there is a lack of unified design principles for these systems: existing Hi-VLA systems differ in how they choose and connect planners, controllers, mechanisms to switch between the two, and how observations and memory are represented in the planner. In this paper, we present a systematic study of Hi-VLA design for robot manipulation. We unify representative Hi-VLA agents under an options-style control framework and benchmark core design choices across short-horizon, long-horizon, and reasoning-intensive tasks. Our analysis distills practical principles for building Hi-VLA systems, showing how model choices and interface mechanisms jointly shape performance. Applying these principles yields a substantially stronger system than either flat VLA control or a naively designed hierarchy, across experiments both in simulation and on a real ALOHA robot. Overall, our results provide a foundation for building more capable, robust, and principled hierarchical VLA agents. More information and video at jiahenghu.github.io/hi-vla.

1 Citations
0 Influential
14.5 Altmetric
73.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!