2605.30117v1 May 28, 2026 cs.AI

VLA-Trace: 표현 및 행동 추적을 통한 비전-언어-행동 모델 진단

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

Xiaozhu Ju
Xiaozhu Ju
Citations: 294
h-index: 7
Y. Zhang
Y. Zhang
Citations: 171
h-index: 7
Yinda Chen
Yinda Chen
Citations: 3
h-index: 1
Jiayu Hu
Jiayu Hu
Citations: 12
h-index: 2
Xiancong Ren
Xiancong Ren
Citations: 11
h-index: 2
Haoyuan Shi
Haoyuan Shi
Citations: 37
h-index: 3
Yingji Zhang
Yingji Zhang
University of Manchester
Citations: 112
h-index: 7
Jinpeng Lu
Jinpeng Lu
Citations: 63
h-index: 4
Qinfang Zhang
Qinfang Zhang
Citations: 87
h-index: 6
Haozhe Shan
Haozhe Shan
Citations: 19
h-index: 2
Hanning Dong
Hanning Dong
Citations: 0
h-index: 0
Yong Dai
Yong Dai
Citations: 56
h-index: 3

비전-언어-행동(VLA) 모델이 다중 모달 지식을 어떻게 신체화된 제어로 변환하는지 이해하는 것은 여전히 해결해야 할 과제입니다. 본 논문에서는 VLA-Trace를 제시합니다. 이는 표현 역학부터 인과적 제어 귀속, 그리고 행동 양상에 이르기까지 통합적인 증거 체인을 통해 VLA 모델을 분석하는 점진적인 진단 프레임워크입니다. 구체적으로, VLA-Trace는 교차 모달 및 체크포인트 드리프트 중심 커널 정렬(CKA)을 사용하여 표현의 변화를 추적하고, 어텐션 제거 개입을 통해 모달리티별 제어 경로를 식별하며, 롤아웃 수준의 행동 탐침을 사용하여 의미론적 연결, 단축 경로 의존성 및 의미 추종을 검사합니다. $π_{0.5}$와 OpenVLA 데이터셋에 대한 실험 결과, 세 가지 주요 사실을 발견했습니다. 첫째, 두 모델은 VLA 미세 조정 과정에서 뚜렷한 모달리티별 적응 동역학을 보입니다. 둘째, 이들은 행동 디코딩 중에 서로 다른 다중 모달 라우팅 전략과 계층별 의존성을 사용합니다. 셋째, VLA 정책은 시각적으로 연계된 경로 생성이 뛰어나지만, 미세한 수준의 의미 추종 능력은 여전히 제한적입니다. 이러한 결과는 표현을 보존하는 적응, 인과적인 VLA 회로 및 조립 가능한 의미론적 제어에 대한 향후 연구 방향을 제시합니다.

Original Abstract

Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic framework that analyzes VLA models through a unified evidence chain from representation dynamics to causal control attribution and behavioral manifestation. It specifically combines cross-modal and checkpoint-drift centered kernel alignment (CKA) to trace representation evolution, attention knockout interventions to identify modality-specific control pathways, and rollout-level behavioral probes to examine grounding, shortcut dependence, and semantic following. Experiments on $π_{0.5}$ and OpenVLA reveal three key findings. First, the two models exhibit distinct modality-specific adaptation dynamics during VLA finetuning. Second, they rely on different multimodal routing strategies and layer-wise dependencies during action decoding. Third, although VLA policies excel at visually grounded trajectory generation, they remain limited in fine-grained semantic following. These findings highlight future directions for representation-preserving adaptation, causal VLA circuits, and compositional semantic control.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!