OmniDrive: LLM 기반 다중 에이전트 월드 모델 - 통합 잠재적 코드 압축을 통한 멀티뷰 주행 비디오 생성
OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation
자율 주행을 위한 생성형 월드 모델은 두 가지 해결되지 않은 문제에 직면합니다. 첫째는 이기종 제어 삽입으로, 자유 형식의 언어, HD 맵, 경로 및 카메라 자세가 호환되지 않는 표현 공간에 존재합니다. 둘째는 사후 융합 방식으로, 각 카메라의 잠재 정보가 전체적인 3차원 기하학적 정보를 제대로 담아내지 못합니다. 우리는 이러한 문제들을 하나의 근본 원인으로 추적했습니다. 즉, 언어, 기하학 및 픽셀을 잠재 토큰 수준에서 연결하는 공유된 상징적 중간 언어의 부재입니다. 본 논문에서는 LLM 기반 다중 에이전트 월드 모델인 DRIVE-CHOREO를 제안합니다. 이 모델은 제어 가능한 멀티뷰 비디오 생성을 잠재 코드 조율로 재구성합니다. 세 개의 Qwen2.5-VL 에이전트 - 사용자 의도를 구조화된 WorldScript로 분석하는 디렉터(Director), 이를 공간적으로 고정된 레이아웃 토큰으로 변환하는 지도 제작자(Cartographer), 그리고 교차 뷰 피드백을 추가적인 감독 신호로 제공하는 감사관(Auditor)이 협력하여 단일한 위치 정보를 포함하는 토큰 시퀀스를 생성합니다. 이 시퀀스는 3D VAE의 컨볼루션 수용 필드 내에서 카메라 간 기하학적 관계를 유지하도록 하는 뷰-타임 순열을 통해 멀티뷰 비디오와 함께 공동으로 압축됩니다. nuScenes 데이터셋에 대한 실험 결과, DRIVE-CHOREO는 멀티뷰 일관성 및 BEV mAP(21.6)에서 최고 성능을 달성했으며, FVD(45.7)도 경쟁력 있는 수준입니다. 또한, 우리 모델로 생성된 합성 데이터를 사용하여 학습한 검출기는 실제 검증 데이터셋에서 +2.4 NDS의 성능 향상을 보여주어 하위 작업에서의 유용성을 입증합니다.
Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion, where per-camera latents fail to encode global 3-D geometry. We trace both to a single root cause: the absence of a shared symbolic interlingua aligning language, geometry, and pixels at the latent-token level. We present DRIVE-CHOREO, an LLM-choreographed multi-agent world model that recasts controllable multi-view video generation as latent choreography. Three Qwen2.5-VL agents - a Director parsing user intent into a structured WorldScript, a Cartographer grounding it into spatially-anchored layout tokens, and an Auditor feeding cross-view critiques back as auxiliary supervision - jointly author a single position-aware token sequence. This sequence is co-compressed with the multi-view video via a view-time permutation that enforces inter-camera geometry within the convolutional receptive field of a 3-D VAE. On nuScenes, DRIVE-CHOREO sets new state-of-the-art multi-view consistency and BEV mAP (21.6) with competitive FVD (45.7); a detector trained purely on our synthetic data gains +2.4 NDS on the real validation split, validating downstream utility.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.