2605.30011v1 May 28, 2026 cs.CV

VisualThink-VLA: 효과적이고 낮은 지연 시간을 갖는 시각-언어-행동 정책을 위한 시각적 중간 추론

VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

Zheqi Lv
Zheqi Lv
Citations: 20
h-index: 2
Wenqiao Zhang
Wenqiao Zhang
Citations: 33
h-index: 3
Yang Dai
Yang Dai
Citations: 18
h-index: 2
Siliang Tang
Siliang Tang
Citations: 1,001
h-index: 13
Ming Gao
Ming Gao
Citations: 25
h-index: 2
Jiaqi Zhu
Jiaqi Zhu
Citations: 1
h-index: 1
Zhiqi Ge
Zhiqi Ge
Citations: 266
h-index: 7
Zixuan Wan
Zixuan Wan
Citations: 3
h-index: 1
Yueting Zhuang
Yueting Zhuang
Citations: 7
h-index: 2
Yuqian Yuan
Yuqian Yuan
Citations: 17
h-index: 3
Binhe Yu
Binhe Yu
Citations: 107
h-index: 2
Haoyuan Zheng
Haoyuan Zheng
Citations: 123
h-index: 7

최근 연구에서는 시각-언어-행동(VLA) 정책에 명시적인 중간 추론 기능을 부여하기 시작했습니다. 그러나, 임베디드 제어 환경에서 텍스트 기반의 사고 과정은 적합하지 않습니다. 관련 없는 또는 약한 의미를 가진 텍스트 정보는 행동 예측을 방해할 수 있으며, 자기 회귀적 텍스트 디코딩은 실시간 폐루프 실행에 너무 많은 지연 시간을 초래합니다. 본 논문에서는 정확하고 낮은 지연 시간을 갖는 VLA 정책을 위한 시각적 중간 추론 프레임워크인 VISUALTHINK-VLA를 제안합니다. VISUALTHINK-VLA의 핵심 철학은 효과적인 시각적 사고를 통해 행동을 유도하는 것입니다. VISUALTHINK-VLA는 공간 정확성을 유지하면서 디코딩 오버헤드를 피하는, 작고 효율적인 시각 정보 인터페이스를 사용하여 행동 예측을 수행합니다. 또한, 성능과 효율성을 더욱 향상시키기 위해, VISUALTHINK-VLA는 시각 정보를 선택적으로 활용하는 맞춤형 라우팅 메커니즘을 채택하여, 시각 정보 토큰을 학습하고 낮은 지연 시간으로 추론을 가능하게 하면서도 높은 용량의 전문성을 유지합니다. 더불어, 우리는 VisualEvidence-Agent를 중심으로 754.7k 개의 VLA 명령어로 구성된 VisualEvidence-Set을 구축하는 감독 및 감사 리소스인 VisualEvidence-Kit을 소개합니다. 여러 벤치마크와 실제 로봇 평가에서 VISUALTHINK-VLA는 대부분의 벤치마크에서 가장 높은 성공률을 달성했으며, 추론 기능을 추가한 기존 모델의 수 초에 달하는 지연 시간을 서브초 수준으로 줄였습니다. 예를 들어, BridgeData V2 데이터셋에서 ECoT 모델의 스텝 지연 시간(8.377초)을 VISUALTHINK-VLA가 0.367초로 단축하여 22.8배의 속도 향상을 보였습니다.

Original Abstract

Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irrelevant or weakly textual information can interfere with action prediction, while autoregressive text decoding adds too much latency for real-time closed-loop execution. We present VISUALTHINK-VLA, a visual intermediate-reasoning framework for accurate, low-latency VLA policies. Our bootstrapping philosophy is to guide action with effective visual thinking: VISUALTHINK-VLA bootstraps action prediction through a compact visual-evidence interface that preserves spatial precision while avoiding decoding overhead. Besides, to further improve performance and efficiency, VISUALTHINK-VLA adopts a tailored selective routing mechanism to learn the visual evidence tokens, enabling low-latency inference while preserving high-capacity specialization. We also introduce VisualEvidence-Kit, a supervision-and-audit resource centered on a VisualEvidence-Agent that constructs a 754.7k VLA instructions VisualEvidence-Set for route supervision and counterfactual faithfulness tests. Across multiple benchmarks and real-robot evaluation, VISUALTHINK-VLA achieves the highest success rate on most benchmarks while reducing the multi-second latency of reasoning-augmented baselines to the sub-second regime. For example, on BridgeData V2, it reduces step latency from 8.377,s with ECoT to 0.367,s, achieving a 22.8 times speedup.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!