운동학 기반 경로 설정, 관찰을 통한 실행: MoE 확장 VLA에서 운동학으로 감독되는 전문가 라우팅
Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA
MoE(Mixture of Experts)는 전문가의 특화 설계를 통해 VLA(Vision-Language Agent)를 향상시키지만, 조작 작업 간의 운동학적 이질성 때문에 라우터가 효과적인 전문가 라우팅을 수행하지 못하는 문제가 있습니다. 더욱이 추론 시에는 운동학적 정보 자체가 제공되지 않는 경우가 많습니다. 본 연구에서는 대부분의 의미적으로 구별되는 조작 작업들이 여러 가지 운동학적 유형으로 축소될 수 있다는 점을 관찰했습니다. 이러한 발견에 따라, 우리는 암묵적인, 관찰 기반 전문가 라우팅에서 벗어나 명시적인, 운동학 기반 전문가 할당 방식으로 전환하는 새로운 패러다임인 '운동학 감독 명시적 라우팅(KinRT)'을 제안합니다. 구체적으로, 우리는 동작 궤적에 대한 운동학 클러스터링을 수행하여 여러 개의 운동학적으로 일관된 그룹으로 나눕니다. 이러한 그룹의 ID는 라우터 학습을 위한 정답 데이터로 사용됩니다. 추론 시에는 라우터가 시각-언어 관찰 정보만을 사용하여 전문가를 할당하며, 동작 운동학에 의존하지 않습니다. KinRT는 실제로 비대칭적인 연결 메커니즘을 도입하여 학습 단계에서 동작 공간의 작업 운동학 정보를 추론 단계에서 관찰 공간으로 전달합니다. 또한, KinRT의 플랫폼 간 일반화 성능을 평가하기 위해 3D 프린팅 기술을 사용하여 경제적이고 DIY(Do-It-Yourself) 로봇 플랫폼(DIYRobot)을 자체적으로 구축했습니다 (2,000 USD 미만). 광범위한 실험 결과는 KinRT가 RoboTwin 벤치마크에서 23.26%, 그리고 우리가 개발한 DIYRobot 플랫폼에서 20.27% 더 우수한 성능을 보인다는 것을 보여줍니다. 저희의 코드와 DIYRobot 플랫폼은 오픈 소스로 공개될 예정입니다.
While MoE augments VLA via expert specialization, router suffers from ineffective expert routing owing to the kinematic heterogeneity of actions across manipulation tasks and, even worse, the unavailability of the kinematic signals at inference time. In this work, we first observe that most semantically distinct manipulation tasks reduce to multiple kinematic archetypes. Motivated by this finding, we propose Kinematics-supervised explicit routing (KinRT), a new paradigm that shifts from implicit, observation-driven expert routing to explicit, kinematics-guided expert dispatching. Specifically, we perform kinematic clustering on action trajectories into multiple kinematically coherent groups, whose IDs serve as ground truth to supervise the training of the router; at inference time, the router dispatches experts only using visual-language observations, without any reliance on action kinematics. KinRT actually introduces an asymmetric bridging mechanism that distills the task kinematics from the action space in training into the observation space at inference. In addition, to assess KinRT's cross-platform generalization, we build an economical, Do-It-Yourself robot (DIYRobot) platform from scratch using 3D-print technology ($<$ 2,000USD). Extensive experiments demonstrate KinRT's superiority over both dense and MoE-featured VLAs by more than 23.26% on RoboTwin benchmark and 20.27% on our introduced DIYRobot platform. Our code and DIYRobot platform will be open-sourced.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.