2608.01755v1 Aug 03, 2026 cs.AI

자율 주행 VLM에서 검증 가능한 추론을 위한 미래 경로의 지연 공개

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Y. Ban
Y. Ban
Citations: 131
h-index: 6

최근 자율 주행(AD)을 위한 비전-언어-액션(VLA) 모델은 시각적 정보와 언어를 결합하여 의사 결정 능력을 향상시키기 위해 체인 오브 소트(Chain-of-Thought, CoT) 방법을 점점 더 많이 사용합니다. 그러나 기존의 어노테이션 파이프라인에서는 일반적으로 지도 모델(teacher model)이 기록된 실제 주행 경로(ground-truth trajectory)를 학습 데이터로 사용합니다. 본 연구는 이를 통해 발생할 수 있는 '경로 고정 편향'(trajectory anchoring bias)을 경험적으로 입증했습니다. 즉, 지도 모델은 제시된 결과에 맞춰 설명을 생성하는 경향이 있으며, 이는 인과 관계에 대한 정확한 추론을 방해하고 환각 현상을 심화시킵니다. 실제 주행 경로 정보를 제거하면 이러한 문제를 완화할 수 있지만, 자유로운 경로 생성은 고수준의 의사 결정과 정밀한 기하학적 합성 및 저수준 동역학 간의 복잡성을 야기합니다. 따라서 본 연구는 경로 수준에서의 자율 주행 의사 결정을 검증하기 위해, 자유로운 경로 생성을 요구하지 않는 새로운 방법인 '자율 주행 객관식 질문(Autonomous-Driving Multiple-Choice Question, AD-MCQ)'을 제안합니다. AD-MCQ는 계획 수립을 명시적인 경로 후보 집합 중에서 선택하는 문제로 재구성합니다. 더 나아가, 본 연구는 미래 경로를 의사 결정 전에 '고정점'으로 사용하는 것이 아니라, 의사 결정 후에 '검증 대상'으로 활용하는 '미래 경로의 지연 공개(Deferred Exposure of Future Trajectories for RLVR, DEFT-RLVR)' 방법을 제안합니다. 실험 결과는 DEFT-RLVR이 자율 주행 추론 능력을 향상시키면서 일반적인 시각적 능력은 유지하거나 오히려 강화함을 보여줍니다. AD-MCQ는 VLM만으로 추론이 가능하며, 후보 경로 생성 방식을 통해 난이도를 조절할 수 있으므로, 검증 가능한 자율 주행 연구를 위한 유연하고 확장 가능한 기반을 제공합니다.

Original Abstract

Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!