2512.07661v4 Dec 08, 2025 cs.CV

최적화 기반 확산 모델을 활용한 인터랙티브 장면 생성

Optimization-Guided Diffusion for Interactive Scene Generation

Tuo An
Tuo An
Citations: 67
h-index: 4
Kashyap Chitta
Kashyap Chitta
University of Tübingen
Citations: 5,711
h-index: 26
Shihao Li
Shihao Li
Citations: 35
h-index: 3
Naisheng Ye
Naisheng Ye
Citations: 12
h-index: 1
Tianyu Li
Tianyu Li
Fudan University
Citations: 1,317
h-index: 14
Peng Su
Peng Su
Citations: 23
h-index: 2
Boyang Wang
Boyang Wang
Citations: 2
h-index: 1
Haiou Liu
Haiou Liu
Citations: 7
h-index: 1
Chenxu Lv
Chenxu Lv
Citations: 997
h-index: 1
Hongyang Li
Hongyang Li
Citations: 773
h-index: 6

자율 주행 차량의 성능 평가를 위해서는 사실적이고 다양한 다중 에이전트 주행 시나리오가 필수적이지만, 이러한 평가에 중요한 안전 관련 이벤트는 드물고 기존 주행 데이터셋에는 충분히 포함되어 있지 않습니다. 데이터 기반 장면 생성은 기존 주행 로그로부터 복잡한 교통 현상을 합성하여 비용 효율적인 대안을 제공합니다. 그러나 기존 모델들은 종종 제어 가능성이 부족하거나 물리적 또는 사회적 제약을 위반하는 결과를 생성하여 활용도가 제한됩니다. 본 논문에서는 구조적 일관성과 상호 작용 인지 능력을 확산 기반 샘플링 과정에서 강화하는 최적화 기반, 학습이 필요 없는 프레임워크인 OMEGA를 제시합니다. OMEGA는 각 역방향 확산 단계를 제약 조건 최적화를 통해 재정렬하여 물리적으로 타당하고 행동적으로 일관된 경로 생성을 유도합니다. 본 프레임워크를 바탕으로, 우리는 에고 차량과 공격자 간의 상호 작용을 게임 이론 기반 최적화 문제로 정의하고, 내쉬 균형을 근사하여 현실적이고 안전 관련성이 높은 적대적인 시나리오를 생성합니다. nuPlan 및 Waymo 데이터셋에 대한 실험 결과, OMEGA는 장면 생성의 현실성, 일관성, 제어 가능성을 향상시키며, 자유로운 탐색 능력을 갖춘 시나리오에서 물리적 및 행동적으로 유효한 장면 비율을 32.35%에서 72.27%로, 제어에 초점을 맞춘 생성에서는 11%에서 80%로 증가시켰습니다. 또한 OMEGA는 전체적인 장면 현실성을 유지하면서 충돌 직전 프레임을 5배 더 많이 생성할 수 있으며, 충돌 시간은 3초 미만입니다.

Original Abstract

Realistic and diverse multi-agent driving scenes are crucial for evaluating autonomous vehicles, but safety-critical events which are essential for this task are rare and underrepresented in driving datasets. Data-driven scene generation offers a low-cost alternative by synthesizing complex traffic behaviors from existing driving logs. However, existing models often lack controllability or yield samples that violate physical or social constraints, limiting their usability. We present OMEGA, an optimization-guided, training-free framework that enforces structural consistency and interaction awareness during diffusion-based sampling from a scene generation model. OMEGA re-anchors each reverse diffusion step via constrained optimization, steering the generation towards physically plausible and behaviorally coherent trajectories. Building on this framework, we formulate ego-attacker interactions as a game-theoretic optimization in the distribution space, approximating Nash equilibria to generate realistic, safety-critical adversarial scenarios. Experiments on nuPlan and Waymo show that OMEGA improves generation realism, consistency, and controllability, increasing the ratio of physically and behaviorally valid scenes from 32.35% to 72.27% for free exploration capabilities, and from 11% to 80% for controllability-focused generation. Our approach can also generate $5\times$ more near-collision frames with a time-to-collision under three seconds while maintaining the overall scene realism.

3 Citations
0 Influential
13 Altmetric
68.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!