2608.05738v1 Aug 06, 2026 cs.RO

문맥 내 VLA: 인컨텍스트 후속 학습과 에이전트 기반 도구 활용을 통해 시각-언어-행동 모델에 언어 능력을 부여

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

Jiaru Yang
Jiaru Yang
Citations: 0
h-index: 0
Wen Huang
Wen Huang
Citations: 1
h-index: 1
Jiale Zhang
Jiale Zhang
Citations: 2
h-index: 1
Maowei Hu
Maowei Hu
Citations: 0
h-index: 0
Hang Guo
Hang Guo
Citations: 1,269
h-index: 12

시각-언어-행동 (VLA) 모델은 일반적인 조작 작업에서 지배적인 방식으로 사용되지만, 거의 모든 VLA 모델은 행동 복제를 통해 훈련됩니다. 즉, 정책은 정적인 이미지와 고정된 지시에 기반하여 전문가의 동작 패턴을 모방합니다. 자연스러운 해결책은 명시적인 추론을 텍스트 기반의 사고 과정 (Chain-of-Thought, CoT)을 통해 주입하는 것입니다. 본 연구에서는 실험적 및 분석적으로, 자유 형식의 텍스트 기반 CoT가 저수준 제어를 저하시킨다는 것을 보여줍니다. 이러한 추론은 현실과 동떨어져 있으며, 지연 시간으로 인해 폐쇄 루프 타이밍이 깨지고, 가장 중요한 것은 추론 토큰과 동작 토큰이 상반된 목표에 최적화되어 정책이 실제로 행동하기보다는 설명하는 것을 학습하게 됩니다. VLA 모델에게 필요한 것은 언어를 생성하는 능력이 아니라, 현실 기반의 언어를 이해하는 능력입니다. 이러한 문제를 해결하기 위해 본 연구에서는 (i) 인컨텍스트 후속 학습 방법을 제안합니다. 이 방법은 시각적 정보를 구조화된 맥락으로 주입하고, 모델을 오직 동작에 대해서만 지도하며 훈련합니다. 또한 (ii) 에이전트 기반의 도구 활용 인터페이스를 통해 정책이 개방형 어휘 감지기, 단안 심도 정보 및 시각-언어 모델을 활용하여 작업과 관련된 정보를 능동적으로 습득하도록 합니다. 저희의 데이터 생성 엔진은 단일의 템플릿화된 설명을 생성하는 대신, 다양한, 재구성되고, 증거 기반의 공간적 설명을 생성합니다. 이를 통해 정책은 이전에 본 적 없는 언어를 해석하는 방법을 학습할 수 있습니다. RoboCasa-GR1, SimplerEnv 및 LIBERO 시뮬레이션 벤치마크와 함께 8가지 실제 로봇 조작 작업에서, 저희 방법은 동일한 구성 하에 CoT 기반 접근 방식과 비교하여 성능과 효율성 모두에서 일관되게 최첨단 (SOTA) 결과를 달성했습니다.

Original Abstract

Vision-Language-Action (VLA) models have become the dominant recipe for generalist manipulation, yet they are almost universally trained by behavior cloning: a policy imitates expert action chunks conditioned on a static image and a fixed instruction. A natural remedy is to inject explicit reasoning through textual chain-of-thought (CoT). We show, both empirically and analytically, that free-form textual CoT degrades low-level control: the reasoning it produces is ungrounded, its latency breaks closed-loop timing, and, crucially, the reasoning and action tokens are optimized against conflicting objectives so that the policy learns to narrate rather than to act. We argue that what a VLA needs is not the ability to generate language, but the ability to consume grounded language. To this end we introduce \textbf{\ourmethod{}}, a framework that endows a VLA with language competence through (i) in-context post-training, in which perceptual evidence is injected as structured context and the model is supervised only on actions, and (ii) an agentic tool-use interface, in which the policy queries open-vocabulary detectors, monocular depth, and a vision--language model to actively acquire task-relevant information. Rather than emitting a single templated caption, our data engine produces diverse, paraphrased, and evidence-conditioned spatial descriptions, so that the policy learns to interpret language it has never seen verbatim. Across the RoboCasa-GR1, SimplerEnv, and LIBERO simulation benchmarks, together with 8 real-world robot manipulation tasks, our method consistently achieves SOTA results in both performance and efficiency when compared with CoT-based approaches under matched configurations.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!