2606.06219v1 Jun 04, 2026 cs.RO

CLEAR: 인지 및 잠재 평가를 통한 엔드 투 엔드 자율 주행 시스템의 적응형 경로 계획

CLEAR: Cognition and Latent Evaluation for Adaptive Routing in End-to-End Autonomous Driving

Zehong Ke
Zehong Ke
Citations: 41
h-index: 3
Zhiyuan Liu
Zhiyuan Liu
Citations: 29
h-index: 3
Yanbo Jiang
Yanbo Jiang
Citations: 47
h-index: 4
Yining Xing
Yining Xing
Citations: 57
h-index: 2
Wenhao Yu
Wenhao Yu
Citations: 4,085
h-index: 6
Jianqiang Wang
Jianqiang Wang
Citations: 50
h-index: 3

엔드 투 엔드 자율 주행 모델은 종종 다중 모드의 조작 생성과 실시간 추론 제약 사이의 균형을 맞추는 데 어려움을 겪습니다. 확산 모델은 다양한 운전 행동을 성공적으로 학습하지만, 반복적인 노이즈 제거 과정은 안전 관련 시스템에 적용하기에는 허용할 수 없는 지연 시간을 초래합니다. 이러한 문제를 해결하기 위해, 우리는 인지 및 잠재 평가를 통한 적응형 경로 계획 프레임워크인 CLEAR (Cognition and Latent Evaluation for Adaptive Routing)를 제안합니다. CLEAR는 매우 빠른 생성적 계획과 심층적인 의미 추론을 결합하며, Drive-JEPA를 시각 인코더로 사용하고 다단계 노이즈 제거 체인을 VAE 잠재 공간에서의 단일 단계 조건부 드리프트로 대체하여 다양성과 전문성 사이의 균형을 맞추는 조건을 도입합니다. 동시에, 우리는 Qwen~3.5~0.8B 모델을 운전 관련 질의응답 데이터셋으로 완전히 미세 조정하여 장면 인지 정보를 담은 숨겨진 상태를 추출합니다. 이러한 상태는 적응형 스케줄러가 사전 정의된 다양한 방식에서 조건 계수 α와 샘플 수 N을 선택하도록 안내하며, 또한 후보 경로 중에서 최적의 경로를 선택하는 크로스 어텐션 점수를 계산하는 데 사용됩니다. NAVSIM v1 벤치마크에서 CLEAR는 최고 수준인 93.7의 PDMS (Planning and Decision Making Score)를 달성했습니다. 우리의 결과는 고품질, 다중 모드 계획이 밀집된 기하학적 정보나 반복적인 샘플링 없이 효율적으로 실행될 수 있음을 보여줍니다.

Original Abstract

End-to-end autonomous driving models often struggle to balance multi-modal maneuver generation with real-time inference constraints. While diffusion models successfully capture diverse driving behaviors, their iterative denoising process incurs unacceptable latency for safety-critical deployment. To address this, we propose CLEAR (Cognition and Latent Evaluation for Adaptive Routing), a framework that combines ultra-fast generative planning with deep semantic reasoning. CLEAR employs Drive-JEPA as the visual encoder and replaces the multi-step denoising chain with a single-step conditional drift in a VAE latent space, introducing a conditioning coefficient to balance diversity and expert precision. Meanwhile, we fully fine-tune Qwen~3.5~0.8B on driving QA pairs to extract scene-aware hidden states. These states guide both an Adaptive Scheduler, which selects the conditioning coefficient $α$ and sample count $N$ from a discrete set of predefined schemes, and a cross-attention scorer that selects the optimal trajectory from candidates. On the NAVSIM v1 benchmark, CLEAR achieves a state-of-the-art PDMS of 93.7. Our results demonstrate that high-fidelity, multi-modal planning can be executed efficiently without dense geometric annotations or iterative sampling.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!