2604.21241v1 Apr 23, 2026 cs.RO

CorridorVLA: 희소 앵커를 사용한 생성적 액션 헤드의 명시적인 공간 제약

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors

Zhuangzhuang Chen
Zhuangzhuang Chen
shenzhen university
Citations: 394
h-index: 11
Dachong Li
Dachong Li
Citations: 47
h-index: 3
Jin Zhang
Jin Zhang
Citations: 84
h-index: 3
Jianqiang Li
Jianqiang Li
Citations: 62
h-index: 3

비전-언어-액션 (VLA) 모델은 종종 다중 모드 입력을 연속적인 제어로 연결하기 위해 중간 표현을 사용하지만, 공간적 지침은 종종 잠재 특징을 통해 암시적으로 주입됩니다. 본 논문에서는 $CorridorVLA$를 제안합니다. $CorridorVLA$는 희소한 공간 앵커를 점진적인 물리적 변화 (예: $Δ$-위치)로 예측하고, 이를 사용하여 액션 생성의 학습 목표에서 명시적인 허용 영역을 설정합니다. 이러한 앵커는 액션 헤드의 공간적 진화를 안내하는 '통로'를 정의하며, 이 통로 밖의 궤적은 교정 그래디언트를 받고, 접촉 및 실행 과정에서의 미세한 편차는 허용됩니다. 더 어려운 LIBERO-Plus 벤치마크에서, $CorridorVLA$는 SmolVLA 및 GR00T 모델 모두에서 일관된 성능 향상을 보이며, 성공률을 해당 기준 모델에 비해 $3.4$%에서 $12.4$%까지 향상시킵니다. 특히, 저희의 GR00T-Corr 모델은 $83.21$%의 성공률을 달성했습니다. 이러한 결과는 액션과 관련된 물리적 신호가 생성적 액션 정책에 직접적이고 해석 가능한 제약을 제공할 수 있으며, 시각적 또는 잠재적 형태로 인코딩된 공간적 지침을 보완할 수 있음을 시사합니다. 코드 및 관련 정보는 https://github.com/corridorVLA 에서 확인할 수 있습니다.

Original Abstract

Vision--Language--Action (VLA) models often use intermediate representations to connect multimodal inputs with continuous control, yet spatial guidance is often injected implicitly through latent features. We propose $CorridorVLA$, which predicts sparse spatial anchors as incremental physical changes (e.g., $Δ$-positions) and uses them to impose an explicit tolerance region in the training objective for action generation. The anchors define a corridor that guides a flow-matching action head: trajectories whose implied spatial evolution falls outside it receive corrective gradients, while minor deviations from contacts and execution noise are permitted. On the more challenging LIBERO-Plus benchmark, CorridorVLA yields consistent gains across both SmolVLA and GR00T, improving success rate by $3.4\%$--$12.4\%$ over the corresponding baselines; notably, our GR00T-Corr variant reaches a success rate of $83.21\%$. These results indicate that action-aligned physical cues can provide direct and interpretable constraints for generative action policies, complementing spatial guidance encoded in visual or latent forms. Code is available at https://github.com/corridorVLA.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!