VisualThink-VLA: 효과적이고 낮은 지연 시간을 갖는 시각-언어-행동 정책을 위한 시각적 중간 추론
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
최근 연구에서는 시각-언어-행동(VLA) 정책에 명시적인 중간 추론 기능을 부여하기 시작했습니다. 그러나, 임베디드 제어 환경에서 텍스트 기반의 사고 과정은 적합하지 않습니다. 관련 없는 또는 약한 의미를 가진 텍스트 정보는 행동 예측을 방해할 수 있으며, 자기 회귀적 텍스트 디코딩은 실시간 폐루프 실행에 너무 많은 지연 시간을 초래합니다. 본 논문에서는 정확하고 낮은 지연 시간을 갖는 VLA 정책을 위한 시각적 중간 추론 프레임워크인 VISUALTHINK-VLA를 제안합니다. VISUALTHINK-VLA의 핵심 철학은 효과적인 시각적 사고를 통해 행동을 유도하는 것입니다. VISUALTHINK-VLA는 공간 정확성을 유지하면서 디코딩 오버헤드를 피하는, 작고 효율적인 시각 정보 인터페이스를 사용하여 행동 예측을 수행합니다. 또한, 성능과 효율성을 더욱 향상시키기 위해, VISUALTHINK-VLA는 시각 정보를 선택적으로 활용하는 맞춤형 라우팅 메커니즘을 채택하여, 시각 정보 토큰을 학습하고 낮은 지연 시간으로 추론을 가능하게 하면서도 높은 용량의 전문성을 유지합니다. 더불어, 우리는 VisualEvidence-Agent를 중심으로 754.7k 개의 VLA 명령어로 구성된 VisualEvidence-Set을 구축하는 감독 및 감사 리소스인 VisualEvidence-Kit을 소개합니다. 여러 벤치마크와 실제 로봇 평가에서 VISUALTHINK-VLA는 대부분의 벤치마크에서 가장 높은 성공률을 달성했으며, 추론 기능을 추가한 기존 모델의 수 초에 달하는 지연 시간을 서브초 수준으로 줄였습니다. 예를 들어, BridgeData V2 데이터셋에서 ECoT 모델의 스텝 지연 시간(8.377초)을 VISUALTHINK-VLA가 0.367초로 단축하여 22.8배의 속도 향상을 보였습니다.
Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irrelevant or weakly textual information can interfere with action prediction, while autoregressive text decoding adds too much latency for real-time closed-loop execution. We present VISUALTHINK-VLA, a visual intermediate-reasoning framework for accurate, low-latency VLA policies. Our bootstrapping philosophy is to guide action with effective visual thinking: VISUALTHINK-VLA bootstraps action prediction through a compact visual-evidence interface that preserves spatial precision while avoiding decoding overhead. Besides, to further improve performance and efficiency, VISUALTHINK-VLA adopts a tailored selective routing mechanism to learn the visual evidence tokens, enabling low-latency inference while preserving high-capacity specialization. We also introduce VisualEvidence-Kit, a supervision-and-audit resource centered on a VisualEvidence-Agent that constructs a 754.7k VLA instructions VisualEvidence-Set for route supervision and counterfactual faithfulness tests. Across multiple benchmarks and real-robot evaluation, VISUALTHINK-VLA achieves the highest success rate on most benchmarks while reducing the multi-second latency of reasoning-augmented baselines to the sub-second regime. For example, on BridgeData V2, it reduces step latency from 8.377,s with ECoT to 0.367,s, achieving a 22.8 times speedup.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.