로봇 액션 표현을 위한 의미론적 고정화
Semantic Anchoring for Robotic Action Representations
비전-언어-액션(VLA) 모델은 사전 학습된 비전-언어 모델로부터 풍부한 의미론적 표현을 상속받지만, 제한적인 로봇 데모 데이터에 대한 미세 조정은 이러한 구조를 저하시키고 일반화 성능을 저해합니다. 따라서 근본적인 질문이 제기됩니다: 좋은 액션 표현이란 무엇일까요? 거울 뉴런 이론에서 제시된 것처럼 관찰과 실행이 의도 수준의 인코딩을 공유한다는 통찰력을 바탕으로, 로봇의 액션 표현이 사전 학습된 인코더가 캡처한 의미론적 구조를 유지하는지 조사했습니다. 체계적인 분석 결과, 미세 조정 과정에서 이러한 구조가 약화되는 것이 확인되었으며, 그 품질은 작업 성공률 및 일반화 성능과 밀접하게 관련되어 있음을 알 수 있었습니다. 또한, 액션 표현을 의미론적 공간에 고정시키고 표현을 공유된 의미 채널과 개별적인 채널로 분해하는 플러그 앤 플레이 방법을 제안합니다. 여기서 개별적인 채널은 추론 과정에서 모두 제거되므로 배포된 모델 자체는 변경되지 않습니다. 시뮬레이션 및 실제 환경 벤치마크에서 다양한 VLA 아키텍처에 대해 검증한 결과, 제안하는 방법은 실제 환경에서의 동일 분포 작업에서 최대 +18.7%, 그리고 일반화 성능에서 최대 +21.5%의 성능 향상을 보였습니다.
Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. A fundamental question therefore arises: what constitutes a good action representation? Inspired by the mirror neuron theory's insight that observation and execution share an intention-level encoding, we examine whether a robot's action representations preserve the semantic structure captured by pretrained encoders. Systematic probing confirms that this structure erodes during finetuning, and that its quality synchronizes with both task success and out-of-distribution generalization. We further introduce a plug-and-play method that anchors action representations to a semantic manifold while decomposing representations into a shared semantic channel and a private channel, all discarded at inference, leaving the deployed model unchanged. Validated on different VLA backbones across simulation and real-world benchmarks, our method yields up to +18.7% on real-world in-distribution tasks and +21.5% on out-of-distribution generalization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.