VLA-Trace: 표현 및 행동 추적을 통한 비전-언어-행동 모델 진단
VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing
비전-언어-행동(VLA) 모델이 다중 모달 지식을 어떻게 신체화된 제어로 변환하는지 이해하는 것은 여전히 해결해야 할 과제입니다. 본 논문에서는 VLA-Trace를 제시합니다. 이는 표현 역학부터 인과적 제어 귀속, 그리고 행동 양상에 이르기까지 통합적인 증거 체인을 통해 VLA 모델을 분석하는 점진적인 진단 프레임워크입니다. 구체적으로, VLA-Trace는 교차 모달 및 체크포인트 드리프트 중심 커널 정렬(CKA)을 사용하여 표현의 변화를 추적하고, 어텐션 제거 개입을 통해 모달리티별 제어 경로를 식별하며, 롤아웃 수준의 행동 탐침을 사용하여 의미론적 연결, 단축 경로 의존성 및 의미 추종을 검사합니다. $π_{0.5}$와 OpenVLA 데이터셋에 대한 실험 결과, 세 가지 주요 사실을 발견했습니다. 첫째, 두 모델은 VLA 미세 조정 과정에서 뚜렷한 모달리티별 적응 동역학을 보입니다. 둘째, 이들은 행동 디코딩 중에 서로 다른 다중 모달 라우팅 전략과 계층별 의존성을 사용합니다. 셋째, VLA 정책은 시각적으로 연계된 경로 생성이 뛰어나지만, 미세한 수준의 의미 추종 능력은 여전히 제한적입니다. 이러한 결과는 표현을 보존하는 적응, 인과적인 VLA 회로 및 조립 가능한 의미론적 제어에 대한 향후 연구 방향을 제시합니다.
Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic framework that analyzes VLA models through a unified evidence chain from representation dynamics to causal control attribution and behavioral manifestation. It specifically combines cross-modal and checkpoint-drift centered kernel alignment (CKA) to trace representation evolution, attention knockout interventions to identify modality-specific control pathways, and rollout-level behavioral probes to examine grounding, shortcut dependence, and semantic following. Experiments on $π_{0.5}$ and OpenVLA reveal three key findings. First, the two models exhibit distinct modality-specific adaptation dynamics during VLA finetuning. Second, they rely on different multimodal routing strategies and layer-wise dependencies during action decoding. Third, although VLA policies excel at visually grounded trajectory generation, they remain limited in fine-grained semantic following. These findings highlight future directions for representation-preserving adaptation, causal VLA circuits, and compositional semantic control.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.