2608.06008v1 Aug 06, 2026 cs.RO

적응형 WAM: 중간 비디오-확산 특징으로부터 품질 기반 초기 종료 계획

Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Sining Ang
Sining Ang
Citations: 11
h-index: 2

대규모 비디오 확산 모델은 자율 주행에 풍부한 시공간 정보를 제공하지만, 기존의 세계-행동 모델들은 종종 반복적인 미래 비디오 생성 과정의 부담을 안고 있습니다. 우리는 더 기본적인 질문을 던집니다: 신뢰할 수 있는 운전 결정을 내리기 위해 비디오 확산 모델의 얼마나 많은 부분을 실행해야 할까요? 비디오 디노이징 타임스텝과 Diffusion Transformer (DiT) 깊이에 대한 통제된 연구를 통해, 계획 성능은 테스트된 비디오 노이즈 수준에 크게 영향을 받지 않는다는 것을 발견했습니다. 반면, 중간 레이어에서도 강력한 경로를 이미 해독할 수 있습니다. 이러한 관찰을 바탕으로, Wan2.2-5B 모델을 기반으로 한 품질 인지형 멀티-익시트 플래너인 Adaptive-WAM을 소개합니다. 트래jectory 확산 헤드가 선택된 DiT 블록에 연결되어 있으며, 경량화된 트래jectory 품질 평가기가 지금까지 해독된 최상의 경로가 특정 품질 임계값을 충족하면 추론을 종료합니다. 그렇지 않으면 캐시된 히든 상태에서 더 깊은 익시트까지 계산이 계속됩니다. 따라서 제안하는 플래너는 미래 비디오 합성을 위해 필요한 반복적인 클래스-프리 디노이징 루프 및 VAE 디코딩 과정을 피하면서, 트래jectory 품질에 따라 백본 모델의 깊이를 동적으로 할당합니다. NAVSIM 데이터셋에서 Adaptive-WAM은 90.8 PDMS를 달성하며, 고정된 익시트를 사용하는 변형은 64개의 제안으로 92.6 PDMS를 달성했습니다. 또한 NAVSIM v2 데이터셋에서 89.9 EPDMS를 얻어, 비교 대상 전방 시야 비디오 세계 모델 플래너 중에서 가장 우수한 결과를 보였습니다. Adaptive-WAM은 타겟 도메인에 대한 추가 학습 없이 nuScenes 데이터셋으로 이전하여 평균 L2 오차 0.88 m 및 충돌률 0.08%를 달성했습니다. A100 GPU에서 Adaptive routing은 PDMS 성능을 90.62에서 90.79로 향상시키면서, 평균 엔드-투-엔드 계획 지연 시간을 170ms로 줄였습니다. 이는 고정된 블록-15 플래너의 190ms보다 약 10% 낮고, 전체 깊이 플래너의 320ms보다 47% 낮은 수치입니다. 코드 공개 예정입니다.

Original Abstract

Large video diffusion models provide rich spatiotemporal priors for autonomous driving, but existing world-action models often inherit the cost of iterative future-video generation even though deployment only requires an ego trajectory. We ask a more basic question: how much of a video diffusion model must be executed to make a reliable driving decision? Through a controlled study of video denoising timesteps and Diffusion Transformer (DiT) depth, we find that planning performance is largely insensitive to the tested video-noise levels, whereas strong trajectories can already be decoded from intermediate layers. Based on this observation, we introduce Adaptive-WAM, a quality-aware multi-exit planner built on a Wan2.2-5B backbone. Trajectory diffusion heads are attached to selected DiT blocks, and a lightweight trajectory-quality scorer terminates inference once the best trajectory decoded so far satisfies a quality threshold; otherwise, computation continues from the cached hidden state to a deeper exit. The deployed planner therefore avoids the iterative classifier-free denoising loop and VAE decoding required for future-video synthesis, while dynamically allocating backbone depth according to trajectory quality. On NAVSIM, the adaptive single-trajectory planner achieves 90.8 PDMS; a separate fixed-exit variant reaches 92.6 PDMS with 64 proposals. It further obtains 89.9 EPDMS on NAVSIM v2, yielding the best reported results among the compared front-view video world-model planners. Without target-domain fine-tuning, Adaptive-WAM transfers to nuScenes with 0.88 m average L2 error and a 0.08\% collision rate. On an A100, adaptive routing improves PDMS from 90.62 to 90.79 while averaging 170 ms end-to-end planning latency, approximately 10\% below the 190 ms fixed block-15 planner and 47\% below the 320 ms fixed full-depth planner. Code will be released.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!