비전-언어-행동 모델을 위한 헤르미트 곡선 기반 경로 사전 지식
Hermite Curves as Trajectory Priors for Vision-Language-Action Models
최근 로봇 조작을 위한 비전-언어-행동(VLA) 모델의 발전에도 불구하고, 액션 덩어리는 여전히 구조화되지 않은 인터페이스입니다. 기존 연구에서는 각 덩어리를 일반적으로 시간 단계별 제어로 평탄화하여 처리하며, 이는 암묵적인 데이터 학습에 의존하게 됩니다. 이러한 학습은 실제 실행 시 울퉁불퉁한 움직임과 경계 불연속성을 초래합니다. 이러한 한계를 해결하기 위해, 우리는 헤르미트 경로 사전 지식을 도입하여 덩어리 경로를 종단점 위치 및 속도를 정의하는 조각별 3차 헤르미트 곡선으로 매개변수화하고, 이를 통해 명시적으로 부드러움과 연속성을 강제합니다. 우리는 이 고정 연산자를 세 가지 변형을 통해 구현했습니다: (1) 헤르미트 토큰(Hermite Tokens), 이는 경계 변수를 자동 회귀적으로 예측합니다; (2) 헤르미트 스캐폴드(Hermite Scaffold), 이는 깔끔한 동작을 기본 스캐폴드와 잔차로 분해합니다; 그리고 (3) 헤르미트 정규화(Hermite Regularization), 이는 사전 지식을 보조 훈련 목표로서 엄격하게 적용합니다. 시뮬레이션 벤치마크 및 실제 로봇 플랫폼에서, 헤르미트 정규화는 세 가지 변형 중 가장 우수한 성능을 달성했으며, LIBERO 데이터셋에서 π0.5 기준 성공률을 95.9%에서 98.7%로, LIBERO-plus 데이터셋에서 85.7%에서 90.9%로, 그리고 네 가지 실제 로봇 작업에서 63.4%에서 90.0%로 향상시켰습니다. 경로 분석 결과, 명시적으로 구조화된 경로 사전 지식이 실행 시간 제약 조건보다는 학습 유도 편향으로 가장 효과적으로 작용하는 것으로 나타났습니다.
Despite recent progress in Vision-Language-Action (VLA) models for robotic manipulation, the action chunk remains a weakly structured interface. Existing work typically flatten each chunk into per-timestep controls, relying on implicit data learning that manifests as jagged motion and boundary discontinuities during physical execution. To address these limitations, we introduce Hermite trajectory priors, parameterizing the chunk trajectory as a piecewise cubic Hermite curve defined by endpoint positions and velocities to explicitly enforce smoothness and continuity. We instantiate this fixed operator across discrete autoregressive and continuous generative paradigms via three variants: (1) Hermite Tokens, which predict quantized boundary variables autoregressively; (2) Hermite Scaffold, which decomposes clean actions into a base scaffold and residuals; and (3) Hermite Regularization, which applies the prior strictly as an auxiliary training objective. Across simulation benchmarks and real-robot platforms, Hermite Regularization achieves superior performance among these three variants, improving π0.5 baseline success rates from 95.9% to 98.7% on LIBERO, 85.7% to 90.9% on LIBERO-plus, and 63.4% to 90.0% across four real-robot tasks without additional inference overhead. Trajectory analyses reveal that explicitly structuring trajectory priors serves most effectively as a learning inductive bias rather than a runtime constraint.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.