2606.30113v1 Jun 29, 2026 cs.RO

SA-VLA: 상태 정보를 활용한 토큰화기를 통한 시각-언어-행동 모델 성능 향상

SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance

Yao Mu
Yao Mu
Citations: 36
h-index: 4
Chunpu Xu
Chunpu Xu
Citations: 207
h-index: 5
Jiayue Kang
Jiayue Kang
Citations: 0
h-index: 0
Teng-Long Jiang
Teng-Long Jiang
Citations: 2
h-index: 1

이산적인 행동 토큰화는 자기 회귀형 VLA (Vision-Language-Action) 정책을 위한 간결한 인터페이스를 제공하지만, 연속적인 로봇 행동을 이산 코드로부터 정확하게 복원하는 것은 여전히 어려운 과제입니다. 기존의 토크나이저는 일반적으로 각 이산 코드를 고정된 연속적인 행동 프로토타입에 매핑하며, 로봇의 현재 고유 감각 상태를 고려하지 않습니다. 이러한 제한은 특히 조작 작업에서 두드러지는데, 동일한 행동 토큰이라도 서로 다른 관절 구성, 물체 자세 및 접촉 조건 하에서 서로 다른 연속적인 제어가 필요할 수 있습니다. 따라서 본 논문에서는 로봇의 상태 정보를 활용하여 행동 디코딩에 영향을 미치는 SA-VLA라는 새로운 상태 인식 행동 토크나이저를 제안합니다. VQ 기반 행동 토큰화에 대한 두 가지 상태 주입 메커니즘을 연구했습니다. 첫 번째는 상태와 행동 특징 간의 크로스 어텐션이며, 두 번째는 상태 정보를 활용하여 각 행동에 대한 조절 계수를 예측하는 경량화된 상태 어댑터입니다. 이 어댑터 방식은 유한 코드북의 효과적인 지원 범위를 확장하여 각 이산 토큰이 상태 의존적인 연속적인 행동들의 집합을 나타낼 수 있도록 하면서도, 이산적인 행동 모델링의 효율성과 호환성을 유지합니다. SA-VLA는 LLM 기반 VLA 정책에 통합되어 있으며, 모델 인터페이스를 최소한으로 변경하여 자기 회귀적 및 병렬적인 행동 토큰 디코딩을 모두 지원합니다. 12개의 RoboTwin 조작 작업에서 SA-VLA는 평균 성공률을 가장 강력한 기존 토크나이저 기준의 0.29에서 0.56으로 향상시켰습니다. 또한, 세 가지 실제 환경에서의 제로샷 시뮬레이션-실제 실험에서 SA-VLA는 평균 성공률을 가장 강력한 기존 토크나이저 기준의 0.15에서 0.33으로 추가적으로 향상시켰습니다. 이러한 결과는 상태 정보를 활용한 행동 디코딩이 이산적인 VLA 정책에서의 압축 간극을 줄이는 간단하고 효과적인 메커니즘임을 보여줍니다.

Original Abstract

Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actions from discrete codes remains challenging. Existing tokenizers typically map each discrete code to a fixed continuous action prototype, ignoring the robot's current proprioceptive state. This limitation is particularly pronounced in manipulation, where the same action token may require different continuous controls under different joint configurations, object poses, and contact conditions. We therefore propose SA-VLA, a state-aware action tokenizer that conditions action decoding on robot state. We study two state-injection mechanisms for VQ-based action tokenization: cross-attention between state and action features, and a lightweight state adapter that predicts action-wise modulation factors for state-conditioned action modulation and reconstruction. The adapter formulation expands the effective support of a finite codebook by allowing each discrete token to represent a family of state-dependent continuous actions, while preserving the efficiency and compatibility of discrete action modeling. Integrated into an LLM-based VLA policy, SA-VLA supports both autoregressive and parallel action-token decoding with minimal changes to the model interface. On 12 RoboTwin manipulation tasks, SA-VLA improves the average success rate from 0.29 to 0.56 over the strongest tokenizer baseline. In zero-shot sim-to-real experiments on three real-world tasks, it further improves average success from 0.15 to 0.33 over the strongest tokenizer baseline. These results demonstrate that state-conditioned action decoding is a simple and effective mechanism for reducing the compression gap in discrete VLA policies.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!