2606.17536v1 Jun 16, 2026 cs.CV

OmniDrive: LLM 기반 다중 에이전트 월드 모델 - 통합 잠재적 코드 압축을 통한 멀티뷰 주행 비디오 생성

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

Zijie Meng
Zijie Meng
Citations: 115
h-index: 6
Wenhua Nie
Wenhua Nie
Citations: 4
h-index: 1
Yufei Liu
Yufei Liu
Citations: 49
h-index: 3
Chen Ma
Chen Ma
Citations: 29
h-index: 3
Zhiyu Li
Zhiyu Li
Citations: 0
h-index: 0
Jiyuan Liu
Jiyuan Liu
Citations: 0
h-index: 0
Bingcai Wei
Bingcai Wei
Citations: 70
h-index: 5
Shuqin Chen
Shuqin Chen
Citations: 159
h-index: 5
Weichen Xu
Weichen Xu
Citations: 25
h-index: 3
Jiquan Yuan
Jiquan Yuan
Citations: 71
h-index: 4
Miao Zhang
Miao Zhang
Citations: 44
h-index: 3

자율 주행을 위한 생성형 월드 모델은 두 가지 해결되지 않은 문제에 직면합니다. 첫째는 이기종 제어 삽입으로, 자유 형식의 언어, HD 맵, 경로 및 카메라 자세가 호환되지 않는 표현 공간에 존재합니다. 둘째는 사후 융합 방식으로, 각 카메라의 잠재 정보가 전체적인 3차원 기하학적 정보를 제대로 담아내지 못합니다. 우리는 이러한 문제들을 하나의 근본 원인으로 추적했습니다. 즉, 언어, 기하학 및 픽셀을 잠재 토큰 수준에서 연결하는 공유된 상징적 중간 언어의 부재입니다. 본 논문에서는 LLM 기반 다중 에이전트 월드 모델인 DRIVE-CHOREO를 제안합니다. 이 모델은 제어 가능한 멀티뷰 비디오 생성을 잠재 코드 조율로 재구성합니다. 세 개의 Qwen2.5-VL 에이전트 - 사용자 의도를 구조화된 WorldScript로 분석하는 디렉터(Director), 이를 공간적으로 고정된 레이아웃 토큰으로 변환하는 지도 제작자(Cartographer), 그리고 교차 뷰 피드백을 추가적인 감독 신호로 제공하는 감사관(Auditor)이 협력하여 단일한 위치 정보를 포함하는 토큰 시퀀스를 생성합니다. 이 시퀀스는 3D VAE의 컨볼루션 수용 필드 내에서 카메라 간 기하학적 관계를 유지하도록 하는 뷰-타임 순열을 통해 멀티뷰 비디오와 함께 공동으로 압축됩니다. nuScenes 데이터셋에 대한 실험 결과, DRIVE-CHOREO는 멀티뷰 일관성 및 BEV mAP(21.6)에서 최고 성능을 달성했으며, FVD(45.7)도 경쟁력 있는 수준입니다. 또한, 우리 모델로 생성된 합성 데이터를 사용하여 학습한 검출기는 실제 검증 데이터셋에서 +2.4 NDS의 성능 향상을 보여주어 하위 작업에서의 유용성을 입증합니다.

Original Abstract

Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion, where per-camera latents fail to encode global 3-D geometry. We trace both to a single root cause: the absence of a shared symbolic interlingua aligning language, geometry, and pixels at the latent-token level. We present DRIVE-CHOREO, an LLM-choreographed multi-agent world model that recasts controllable multi-view video generation as latent choreography. Three Qwen2.5-VL agents - a Director parsing user intent into a structured WorldScript, a Cartographer grounding it into spatially-anchored layout tokens, and an Auditor feeding cross-view critiques back as auxiliary supervision - jointly author a single position-aware token sequence. This sequence is co-compressed with the multi-view video via a view-time permutation that enforces inter-camera geometry within the convolutional receptive field of a 3-D VAE. On nuScenes, DRIVE-CHOREO sets new state-of-the-art multi-view consistency and BEV mAP (21.6) with competitive FVD (45.7); a detector trained purely on our synthetic data gains +2.4 NDS on the real validation split, validating downstream utility.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!