SUV: 미래 장면 이해를 비디오 생성으로 활용한 엔드 투 엔드 자율 주행
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
엔드 투 엔드 자율 주행은 미래 장면의 일관성 있는 이해를 필요로 하지만, 기존 방법들은 이러한 장면을 작업에 특화된 모듈과 출력 형식으로 모델링하여 확장성이 제한적입니다. 본 연구에서는 비디오 생성이 공유 예측기로 활용될 수 있는지 탐구합니다. 우리는 SUV(Scene Understanding as Video)라는 통합형 엔드 투 엔드 자율 주행 프레임워크를 제안하며, 이는 사전 훈련된 비디오 기반 모델을 사용하여 미래 장면 이해를 비디오 생성 문제로 정의합니다. SUV는 미리 학습된 비디오 전문 모듈을 활용하여 미래의 시각적 표현, 의미론적 정보, 상대적인 깊이 및 객체 수준의 동역학을 비디오 스트림으로 모델링하며, 각 스트림에 특화된 시각 예측 모듈은 사용하지 않습니다. 또한, 비디오-액션 주의 메커니즘을 통해 액션 전문 모듈은 모든 미래 스트림의 잠재 표현에 주목하여 에고 차량의 주행 경로를 생성합니다. 실험 결과, SUV는 네 가지 미래 스트림을 직접적으로 예측하며, 추가적인 분석을 통해 구조화된 미래 정보 및 미래 스트림에 대한 직접적인 접근이 더 높은 주행 경로 계획 성능을 가져옴을 확인했습니다. 단일 전방 카메라만을 사용하고 후보 주행 경로 선택 과정을 거치지 않음에도 불구하고, SUV는 NAVSIM-v2 데이터셋의 navtest와 navhard 두 가지 부분에서 최첨단 방법들을 능가하는 91.0 EPDMS 및 36.9의 성능을 달성했습니다. 또한, long-tail WOD-E2E 벤치마크에서는 7.94의 경쟁력 있는 RFS를 기록했습니다.
End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance-level dynamics as video streams with a shared video expert, without stream-specific visual prediction heads. Through joint video-action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future-stream access yield higher trajectory planning scores. With only a single front camera and no candidate-trajectory selection, SUV outperforms a broad set of recent state-of-the-art methods on both NAVSIM-v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long-tail WOD-E2E benchmark, SUV achieves a competitive RFS of 7.94.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.