2607.29052v1 Jul 31, 2026 cs.RO

결과 지향적 증류: 자율 주행 환경에서 VLM 추론 능력을 향상시키는 교수-학생 프레임워크

Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Yu Wu
Yu Wu
Citations: 13,161
h-index: 27
Zeyu Dong
Zeyu Dong
Citations: 4
h-index: 1
Yimin Zhu
Yimin Zhu
Citations: 19
h-index: 2

종단 간(End-to-End, E2E) 자율 주행 시스템은 시각 정보를 기반으로 직접적인 제어 동작을 학습하는 것을 목표로 합니다. 하지만 이러한 E2E 모델은 종종 블랙박스처럼 작동하며 복잡한 상황에서 어려움을 겪습니다. 이를 해결하기 위해 최근 연구에서는 해석 가능성과 주행 안정성을 향상시키기 위해 비전-언어 모델(Vision-Language Models, VLMs)을 활용하여 명시적인 추론 기능을 추가하고 있습니다. 이러한 접근 방식은 일반적으로 미리 생성된 어노테이션에 의존하는데, 이는 잠재적으로 오류가 있는 레이블을 포함하며 상당한 인적 자원을 필요로 합니다. 본 연구에서는 구조화된 추론과 기하학적 정밀도를 통합하는 교수-학생 아키텍처를 기반으로 하는 새로운 프레임워크를 제안합니다. 교수 모델은 반사적 추론을 도입하여 VLM이 논리적인 설명을 생성하고, 실제 동작에 대한 지도 하에 이러한 설명을 개선하도록 합니다. 이를 통해 중간 레이블 없이도 일반화 능력을 향상시킬 수 있습니다. 학생 모델은 지도 학습을 통해 교수 모델의 추론 능력을 증류합니다. 또한, 텍스트 기반 추론을 연속적인 경로로 변환하는 별도의 웨이포인트 디코더를 설계했습니다. 제안된 솔루션은 명확한 추론을 제공하여 해석 가능성을 높이고, 동시에 안정적이고 정확한 주행 성능을 달성하는 두 가지 목표를 통합합니다. 이는 단계별 추론 엔진 내에서 이러한 두 가지 목표 간의 시너지 효과를 활용하여 주행 성능을 향상시키고, 추론 결과를 사용하여 주행 예측을 안내합니다. Waymo 벤치마크에서 평가한 결과, 제안된 프레임워크는 기존의 추론 기반 모델보다 제로샷 추론 능력, 웨이포인트 정확도 및 추론 효율성 측면에서 더 뛰어난 성능을 보였습니다. 실험 결과를 통해 제안하는 디자인이 검증되었으며, 추론 텍스트가 주행 추론에 상당한 기여를 한다는 것을 확인했습니다. 이는 추론 기능을 사용하지 않는 동일한 모델과 비교하여 약 24%의 성능 향상을 가져왔습니다. 본 연구는 자율 주행 기술을 해석 가능하고 실제 적용 가능한 시스템으로 발전시키는 데 기여합니다.

Original Abstract

End-to-end (E2E) autonomous driving aims to learn a direct mapping from visual observations to control actions. However, these E2E models often act as black boxes and struggle with complex scenarios. To address this, recent works incorporate Vision-Language Models (VLMs) to provide explicit reasoning, enhancing both interpretability and driving robustness. These approaches typically rely on pre-generated annotations, which suffer from potentially flawed labels and require costly human labor. In this work, we propose a new framework that integrates structured reasoning and geometric precision through a teacher-student architecture. The teacher model introduces reflective reasoning, where the VLM generates logical explanations and then reflectively refines the reasoning under the supervision of ground-truth action. This enhances zero-shot generalization without intermediate labels. The student model distills the teacher's reasoning capabilities via supervised fine-tuning. We also design a separate waypoint decoder that interprets textual reasoning into continuous trajectories. Our proposed solution integrates two goals: providing explicit reasoning for interpretability and delivering robust and accurate driving performance. It leverages the synergy between these two objectives within a staged inference engine to enhance driving performance and explicitly uses the reasoning to guide driving prediction. Evaluated on Waymo benchmarks, our framework outperforms classical reasoning-based baselines in zero-shot reasoning, waypoint accuracy, and inference efficiency. Our experiments validate this design, demonstrating that the reasoning text makes a significant contribution to driving inference, resulting in around a 24% improvement in performance compared to an identical model that lacks reasoning. Our work advances reasoning-driven autonomous driving toward interpretable and deployable systems.

0 Citations
0 Influential
13.5 Altmetric
67.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!