2606.05737v1 Jun 04, 2026 cs.CV

단순하게: 비전-언어-액션 모델을 위한 단일 단계 액션 생성

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

Xipeng Qiu
Xipeng Qiu
Citations: 28
h-index: 2
Shiduo Zhang
Shiduo Zhang
Citations: 375
h-index: 8
Jingjing Gong
Jingjing Gong
Citations: 174
h-index: 6
Yitong Chen
Yitong Chen
Citations: 17
h-index: 2

확산 기반 비전-언어-액션(VLA) 모델은 종종 이미지 생성 방식을 따릅니다. 즉, 액션은 반복적인 디노이징 과정을 통해 생성됩니다. 본 논문에서는 VLA 액션 생성이 다른 조건-목표 구조를 가진다는 점을 지적합니다. 정책은 풍부한 관찰 데이터, 언어 정보 및 상태에 의해 결정되지만, 예측하는 것은 작고 저차원의 액션입니다. 이러한 비대칭성을 고려할 때, 이미지 합성에 개발된 고급 단일 단계 방법을 반드시 사용할 필요는 없습니다. 저희는 표준 속도 예측 방식을 사용하고, 교사 모델, 증류 단계 또는 보조 목적 함수를 추가하지 않았습니다. 핵심 아이디어는 단순히 학습 시간 분포를 높은 노이즈 상태로 편향시키는 것입니다. 먼저 제어된 MNIST 그리드-시퀀스 작업에서 이 효과를 확인한 후, 다양한 로봇 제어 실험을 통해 검증했습니다. 표준 LIBERO, LIBERO-Plus 및 LIBERO-Pro 환경에서, 높은 노이즈 편향으로 학습된 단일 단계 정책은 동일한 방식으로 훈련된 10단계 디코딩과 유사한 성능을 보였으며, 특히 표준 LIBERO 환경에서는 균일한 시간 분포로 훈련된 10단계 정책보다 더 나은 결과를 얻었습니다. 실제 로봇을 사용한 양손 YAM RSS 평가를 통해 동일한 샘플링 추세를 확인할 수 있었습니다. 14억 개의 파라미터를 가진 대규모 언어 모델(VLM)과 3천만 개의 액션 헤드를 사용하여, 단일 단계 디코딩은 LIBERO-Long 데이터셋에서 95.6%의 정확도를 달성했습니다. 이러한 결과는 강력한 단일 단계 VLA 액션 생성이 이미지 생성에 사용되는 복잡한 다단계 확산 기법을 도입하지 않고도 표준 확산 학습을 통해 자연스럽게 발생할 수 있음을 보여줍니다.

Original Abstract

Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising. We argue that VLA action generation has a different condition-target structure: the policy is conditioned on rich observations, language, and state, but predicts only a compact, low-dimensional action chunk. Under this asymmetry, strong one-step action generation should not necessarily require the advanced one-step methods developed for image synthesis. We keep standard velocity prediction and add no teacher model, distillation stage, or auxiliary objective; in our main recipe, we simply bias the training time distribution toward high-noise states. We first isolate the effect in a controlled MNIST grid-to-sequence task, then test it with extensive robot-policy experiments. Across standard LIBERO, LIBERO-Plus, and LIBERO-Pro, one-step policies trained with high-noise biased schedules generally match ten-step decoding under the same recipe, and on standard LIBERO can exceed ten-step policies trained with a uniform time distribution. A real-robot bimanual YAM RSS evaluation gives a small-sample cross-architecture check of the same sampler trend. On a 1.4B VLM model with a 30M action head, one-step decoding reaches 95.6\% on LIBERO-Long. These results show that strong one-step VLA action generation can emerge from standard diffusion training, without importing the full few-step diffusion machinery developed for image generation.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!