2606.22913v1 Jun 22, 2026 cs.CV

의도, 숙고, 개선: 자율 주행을 위한 적응형 다중 모드 반사 프레임워크

Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving

Jianhua Han
Jianhua Han
Citations: 2,947
h-index: 29
Xiaodan Liang
Xiaodan Liang
Citations: 221
h-index: 8
Likui Zhang
Likui Zhang
Citations: 10
h-index: 2
Tao Tang
Tao Tang
Citations: 93
h-index: 3
Xiuwei Chen
Xiuwei Chen
Citations: 38
h-index: 3
Zisheng Chen
Zisheng Chen
Citations: 38
h-index: 3
Yuping Qiu
Yuping Qiu
Citations: 0
h-index: 0
Ying-Cong Chen
Ying-Cong Chen
Citations: 40
h-index: 3
Hang Xu
Hang Xu
Citations: 280
h-index: 8

최근의 비전-언어-행동(VLA) 모델은 추론 기능을 통합하여 해석 가능성과 계획 품질을 향상시켜 자율 주행 분야에 큰 발전을 가져왔습니다. 그러나 대부분의 기존 접근 방식은 최종 경로를 직접 생성할 뿐, 그 미래 결과에 대해 명시적으로 검토하지 않기 때문에 복잡하고 역동적인 환경에서 신뢰성이 제한됩니다. 이러한 한계를 해결하기 위해, 우리는 자율 주행을 위한 적응형 다중 모드 반사 프레임워크인 IRR-Drive (Intend, Reflect, Refine)를 제안합니다. 구체적으로, IRR-Drive는 고차원적인 추론과 물리적 제약 조건을 긴밀하게 결합하기 위해 먼저 초기 의도를 텍스트 형태로 생성하고, 미래의 잠재적인 상호 작용을 예측하여 미래의 의미론적 탑-다운(bird's-eye view, BEV) 표현을 예상합니다. 이러한 이중 모드 (텍스트 + BEV) 반사 공간은 예상되는 장면 변화를 명시적으로 모델링하여, 모델이 최종 경로를 생성하기 전에 초기 의도를 엄격하게 수정하고 개선할 수 있도록 합니다. 또한, 계획 성능과 계산 효율성의 균형을 맞추기 위해, 우리는 반사를 위한 학습 데이터를 구축하고 적응적인 반사 보상을 설계했습니다. 이를 통해 모델은 장면의 복잡성에 따라 추론 모드를 적응적으로 선택할 수 있습니다. IRR-Drive는 추론을 단순한 해석 보조 기능으로 사용하는 대신, 적응적인 반사 메커니즘을 계획 프레임워크에 직접 통합하여, 장면 복잡성에 의해 구동되는, 현실 기반의 의사 결정을 고려한 경로 수정 기능을 가능하게 합니다. 우리의 방법은 NAVSIM 벤치마크에서 PDMS 및 EPDMS 모두에서 최첨단 성능을 달성했습니다. 광범위한 실험 결과는 다중 모드 반사 프레임워크의 효과를 입증하고 제안된 적응적인 반사 전략의 유효성을 검증합니다.

Original Abstract

Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planning quality. However, most existing approaches directly generate the final trajectory without explicitly examining its future consequences, which limits their reliability in complex and dynamic environments. To address this limitation, we propose IRR-Drive (Intend, Reflect, Refine), an adaptive multimodal reflection framework for autonomous driving. Specifically, to tightly couple high-level reasoning with physical constraints, IRR-Drive first generates a preliminary textual intention and anticipates potential interactions by predicting future semantic bird's-eye view (BEV) representations. This dual-modality (Text + BEV) reflection space explicitly models anticipated scene evolution, enabling the model to rigorously self-correct and refine its initial intent before generating the final trajectory. Furthermore, to balance planning performance and computational efficiency, we construct reflection-oriented training data and design an adaptive reflection reward, enabling the model to adaptively select its reasoning mode according to scene complexity. Instead of using reasoning primarily as an auxiliary interpretation, IRR-Drive directly integrates an adaptive reflection mechanism into the planning framework, enabling grounded, decision-aware trajectory correction that is driven by scene complexity. Our method achieves state-of-the-art performance on the NAVSIM benchmark in both PDMS and EPDMS. Extensive experiments demonstrate the effectiveness of our multimodal reflection framework and validate the efficacy of the proposed adaptive reflection strategy.

0 Citations
0 Influential
14.5 Altmetric
72.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!