VLA 모델의 테스트 시간 모달 적응을 위한 원인 관계 기반 추론-진단-수정 프레임워크
A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models
비전-언어-액션(VLA) 모델은 시각적 관찰과 고유 체감 정보를 바탕으로 언어 지침에 따라 작업을 수행하기 위한 순차적인 액션을 예측합니다. 그러나 로봇 조작에는 원거리 이동 및 근거리 상호 작용과 같이 시간이 지남에 따라 시각적 관찰의 중요도가 변하는 동적 단계가 포함되므로, VLA 모델에서 모달을 어떻게 결합할 것인지는 여전히 해결해야 할 문제입니다. 본 논문에서는 다양한 VLA 아키텍처와 통합하여 테스트 시간에 액션 예측을 개선할 수 있는 모델에 국한되지 않는 추론-진단-수정(IDR) 프레임워크를 제안합니다. IDR은 먼저 시각적 관찰의 사실적인 및 반사실적인 시나리오에서 액션을 추론하고, 추정된 동적 중요도로 시각적 관찰의 인과적 효과를 진단한 후, 학습 없이 액션 예측을 개선하는 데 사용합니다. 또한, 본 논문에서는 IDR 프레임워크를 구현하기 위해 원인 관계를 고려한 액션 리파이너를 설계하였으며, 여기에는 반사실적인 액션을 추론하기 위한 제로 패딩 개입, 인과적 효과를 진단하기 위한 정규화 기반 양자화, 그리고 액션을 개선하기 위한 게이티드 잔차 결합이 포함됩니다. 시뮬레이션 벤치마크 및 실제 작업에서의 광범위한 실험 결과는 다양한 VLA 모델에서 전반적인 성능 향상을 보여주며, 테스트 시간에 시각적 중요도를 동적으로 조정하는 것의 효과를 입증합니다.
Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructions, conditioned on visual observations and proprioceptive states. However, how to fuse modalities in VLA models remains an open problem, since robot manipulation involves dynamic phases, such as long-distance movements and close-range interactions, in which the importance of visual observations may vary over time. In this paper, we propose an infer-diagnose-refine (IDR) framework, a model-agnostic framework that can be integrated with diverse VLA architectures for refining action predictions at test time. IDR first infers actions under factual and counterfactual scenarios of visual observations, and then diagnoses the causal effects of visual observations as the estimated dynamic importance, which is finally used to refine the action predictions in a training-free manner. We further design a causality-aware action refiner to realize the IDR framework, including zero-padding interventions for inferring counterfactual actions, norm-based quantification for diagnosing causal effects, and gated residual fusion for refining actions. Extensive experiments on both simulation benchmarks and real-world tasks show improvements in overall performance across multiple VLA backbones, demonstrating the efficacy of dynamically adjusting visual importance at test time.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.