2607.26513v1 Jul 29, 2026 cs.RO

시각-언어-행동 모델을 위한 명시적인 운동 제어: 분석적 개념 기반 접근 방식

Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models

Mingyang Sun
Mingyang Sun
Citations: 136
h-index: 5
Jiude Wei
Jiude Wei
Citations: 18
h-index: 3
Xiujian Liang
Xiujian Liang
Citations: 0
h-index: 0
Qichen He
Qichen He
Citations: 0
h-index: 0
Donglin Wang
Donglin Wang
Citations: 11
h-index: 2
Cewu Lu
Cewu Lu
Citations: 2,251
h-index: 22
Jianhua Sun
Jianhua Sun
Citations: 670
h-index: 8

현재의 시각-언어-행동 (VLA) 모델은 주로 2차원 입력에 의존하며, 3차원 물리 세계에 내재된 풍부한 객체 구조 정보와 상식 지식을 간과합니다. 이러한 한계는 VLA 모델의 공간 인지 능력과 복잡하고 정밀한 조작을 위한 적응성을 제한합니다. 이 중요한 격차를 해소하기 위해, 우리는 VLA 모델에서 실행 가능한 분석적 개념을 구축하는 Concept Expert 모듈을 개발했습니다. 이 모듈은 객체를 명시적인 프로그래밍 청사진으로 표현합니다. 우리의 메커니즘은 두 가지 상호 보완적인 단계로 작동합니다. 첫째, VLA 추론 전에 Concept Expert는 시각 기반 모델 (VFM)로부터 3차원 정보를 활용하여 초기 운동 및 구조 파라미터를 추정합니다. 둘째, 조작 과정 전반에 걸쳐 VLA 모델은 자체적으로 동적 개념 파라미터를 추적하고, 관찰된 변화와 지속적으로 일치시켜 정확성을 유지합니다. 구축된 분석적 개념은 (1) 밀집되고 프로그래밍 방식의 조작 보상과 (2) 정밀한 공간 안내를 통해 VLA 미세 조정에 명시적이고 고품질의 지침을 제공합니다. 이러한 접근 방식을 통해 VLA 모델은 물리적으로 기반한 상호 작용 동작을 학습하면서도 엔드 투 엔드 학습의 유연성을 유지할 수 있습니다. 우리의 실험 결과는 지도 및 강화 학습 환경 모두에서 성공률과 학습 효율성 측면에서 일관된 개선을 보여주며, 구조화되고 개념 기반의 지침이 VLA 모델의 후속 훈련에 효과적임을 입증합니다.

Original Abstract

Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inherent in the 3D physical world. This deficiency restricts their spatial awareness and adaptability for complex, high-precision manipulation. To bridge this crucial gap, we construct a Concept Expert module for VLA to build executable Analytic Concepts that represent objects as explicit, programmatic blueprints. Our mechanism operates in two synergistic phases: First, prior to VLA inference, the Concept Expert leverages 3D information from Vision Foundation Models (VFMs) to estimate the initial kinematic and structural parameters. Second, throughout the manipulation process, the VLA model utilizes its inherent capability to dynamically track the dynamic concept parameters, continuously aligning them with observational changes to ensure persistent accuracy. Once established, the Analytic Concepts provide explicit, high-quality guidance for VLA fine-tuning through (1) dense, programmatic manipulation rewards and (2) precise spatial guidance. This formulation allows VLA models to learn physically grounded interaction behaviors while maintaining end-to-end learning flexibility. Our experimental results show consistent improvements in success rate and learning efficiency across supervised and reinforcement learning settings, demonstrating the effectiveness of structured, concept-based guidance for VLA post-training.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!