2607.04681v1 Jul 06, 2026 cs.RO

비전-언어-행동 모델은 정말 의미하는 바를 전달할까? 몸체화된 추론에서 신뢰성의 역할에 대한 연구

Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning

Matthew Foutter
Matthew Foutter
Citations: 143
h-index: 4
Matteo Cercola
Matteo Cercola
Citations: 5
h-index: 1
Lena Wild
Lena Wild
Citations: 23
h-index: 2
Yunshan Wang
Yunshan Wang
Citations: 0
h-index: 0
Michelle Li
Michelle Li
Citations: 0
h-index: 0
Daniele Gammelli
Daniele Gammelli
Citations: 462
h-index: 11
Marco Pavone
Marco Pavone
Citations: 2,785
h-index: 26

몸체화된 연쇄적 사고(Embodied Chain-of-Thought)는 블랙박스 비전-언어-행동(VLA) 모델의 로봇 의사 결정 및 해석력을 향상시키는 유망한 메커니즘으로 떠오르고 있습니다. 하지만, 이러한 방식으로 표현되는 연쇄적 사고가 실제 정책의 근본적인 의사 결정 과정을 진실되게 반영하는지에 대한 이해는 아직 부족합니다. 본 연구에서는 작업 성능을 향상시키는 기능적 추론과, 정책의 내부 의사 결정 과정을 실제로 반영하는 신뢰성 있는 추론을 구분합니다. 우리는 최첨단(SoTA) 정렬 전략이 필요한 부분이지만 충분하지 않은 수준의 신뢰성을 제공하며, 중간 단계에서 원인-결과 관계를 가리고 정책 일반화를 제한할 수 있는 추론 (예: 환경에 기반하지 않거나 내부적으로 단절되거나 일관성이 없는 추론)을 허용한다고 주장합니다. 우리는 인간 평가를 통해 자율 주행을 위한 최첨단 추론 모델의 사례를 분석하여, 추론 품질과 하위 작업 성능 향상 간의 불일치를 보여줍니다. 그런 다음, 학습된 비평기(critic), Pinocchio를 사용하여 몸체화된 신뢰성을 측정하는 행동 기반 지표를 정의하고, 이를 통해 관찰 데이터의 적합성과 단계별 일관성을 평가하여 보상을 제공합니다. 강화 학습을 통해 훈련된 본 모델은 다양한 자율 주행 테스트 환경에서 최첨단 정렬 및 경로 오류 최소화를 목표로 한 기존 방식보다 신뢰성을 각각 4% 및 18% 향상시키면서도 경쟁력 있는 하위 작업 성능을 유지합니다. 마지막으로, 합성된 데이터셋을 활용한 테스트 결과, 신뢰성 향상을 위한 후속 학습은 최첨단 정책에 비해 희귀한 반사실적 시나리오에 대한 정책의 반응성을 최대 1.6배 향상시켜, 신뢰성 있는 추론이 더욱 강력하고 일반화 가능하며 해석 가능한 몸체화된 지능을 구현하는 데 기여한다는 것을 시사합니다. 프로젝트 페이지: https://mjf-su.github.io/pinocchio/

Original Abstract

Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models. However, whether this verbalized Chain-of-Thought truthfully reflects the policy's underlying decision process remains poorly understood. We distinguish between functional reasoning, in which reasoning improves task performance, and faithful reasoning, in which reasoning truly reflects the policy's internal decision process. We argue that SoTA alignment strategies offer a necessary but insufficient notion of faithfulness, admitting reasoning whose intermediate steps can mask the causal links in action prediction through confounding factors (e.g., reasoning that is ungrounded in the environment and internally disconnected or inconsistent), restricting policy generalization. We study this gap through a human evaluation of a SoTA reasoning model for autonomous driving, revealing an inconsistent coupling between reasoning quality and downstream trajectory improvement. We then operationalize a behavioral surrogate for embodied faithfulness through a learned critic, Pinocchio, scoring observation grounding and stepwise coherence, and use this critic as a dense reward signal in post-training an embodied policy with reinforcement learning. Across withheld driving benchmarks, our post-trained planner improves faithfulness by 4% and 18% over SoTA alignment and trajectory error post-training baselines, respectively, while maintaining competitive downstream task performance. Finally, on a synthetic out-of-distribution test set, post-training for faithfulness improves policy responsiveness to rare counterfactual scenarios by 1.6x that of a SoTA policy, suggesting that faithful reasoning traces contribute to more robust, generalizable, and interpretable embodied intelligence. Project page: https://mjf-su.github.io/pinocchio/

0 Citations
0 Influential
13 Altmetric
65.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!