2608.04765v1 Aug 05, 2026 cs.RO

시각-언어-행동 모델에서 장기 계획을 위한 명시적 언어 메모리

Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Ziyi Ye
Ziyi Ye
Citations: 765
h-index: 8

시각-언어-행동(VLA) 모델은 시각 인지, 언어 이해 및 로봇 제어를 통합하는 중요한 패러다임을 제공합니다. 그러나 기존 VLA 모델은 여전히 장기 작업에서 심각한 어려움에 직면하고 있습니다. 제한적인 전문가 데모는 작업 간의 조합 일반화를 저해하며, 장기 작업의 비마르코프 특성은 현재 관찰 정보만으로 정책을 수립하는 경우 시간적 일관성을 유지하기 어렵게 만듭니다. 또한, 제한된 폐루프 오류 수정 기능은 실행 중 발생하는 오류를 누적시키며, 엔드-투-엔드 행동 미세 조정은 시각-언어 모델(VLM)의 핵심적인 의미 표현을 약화시킬 수 있습니다. 이러한 문제점을 해결하기 위해, 우리는 명시적인 언어 메모리 모듈을 갖춘 계층적 장기 VLA 아키텍처를 제안합니다. 핵심 아이디어는 이산적인 시간 관찰 정보를 시간 논리를 활용하여 일관된 텍스트 메모리 시퀀스로 변환하는 것입니다. 이 시스템은 고수준 VLM과 저수준 VLA로 분리됩니다. 고수준 VLM은 시각 질의 응답 학습 패러다임을 통해 의미 추론을 수행하며, 저수준 VLA는 하위 작업 지침 및 시각 관찰 정보에 기반하여 정밀한 연속 제어를 실행합니다. 고수준 VLM은 이전 메모리를 문맥 앵커로 사용하여 언어 메모리와 하위 작업 지침을 반복적으로 업데이트함으로써 장기 실행 중에 지속적인 시간 추적과 동적인 오류 수정을 가능하게 합니다. 우리는 제안된 방법을 여러 시뮬레이션 환경에서 평가하고 실제 로봇 플랫폼에서의 시뮬레이션-실제(sim-to-real) 실험을 수행했습니다. 결과는 명시적인 언어 메모리가 복잡한 장기 작업에서 VLA 모델의 성공률과 안정성을 향상시키고 의사 결정 과정을 해석 가능하게 만든다는 것을 보여줍니다.

Original Abstract

Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in long-horizon tasks: sparse expert demonstrations constrain cross-task compositional generalization; the non-Markovian nature of long-horizon tasks makes it difficult for policies conditioned only on current observations to maintain temporal consistency; limited closed-loop error correction allows execution errors to accumulate; and end-to-end action fine-tuning may weaken the high-level semantic representations of vision-language model (VLM) backbones. To address these issues, we propose a hierarchical long-horizon VLA architecture with an explicit language-memory module. The central idea is to convert discrete temporal observations into a coherent textual memory sequence with temporal logic. The system is decoupled into a high-level VLM and a low-level VLA: the high-level VLM performs semantic reasoning through a visual question answering training paradigm, while the low-level VLA executes precise continuous control conditioned on subtask instructions and visual observations. The high-level VLM recursively updates both language memory and subtask instructions using the previous memory as a contextual anchor, enabling persistent temporal tracking and dynamic correction during long-horizon execution. We evaluate the proposed method in multiple simulation environments and conduct sim-to-real experiments on a real robotic platform. The results demonstrate that explicit language memory improves the success rate and robustness of VLA models on complex long-horizon tasks while providing an interpretable semantic account of the decision process.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!