CARLA-GS: 자율 주행의 특이 상황 생성에 필요한 시각적 표현, 추론 및 물리 시뮬레이션 분리
CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis
자율 주행 시스템의 안전성 평가는 드물지만 매우 중요한 안전 관련 상호작용에 의해 결정되므로, 실제와 유사한 이미지를 통해 의도적으로 특이 상황을 생성할 수 있는 시뮬레이터가 필요합니다. 특이 상황 생성은 시각적 표현, 장면 추론 및 차량 경로 생성 및 제어 등 다양한 요소를 포괄하는 복합적인 문제입니다. 기존의 지식 기반 및 모델 기반 접근 방식은 일반적으로 장면 또는 경로 구성 요소에 개별적으로 집중하는 반면, 확산 모델 기반 방법은 전체 과정을 시도하지만 여전히 공간-시간 일관성과 물리적 현실성을 확보하는 데 어려움을 겪습니다. 이러한 다양한 측면을 하나의 프레임워크로 통합하기 위해, 우리는 시각적 표현, 의미론적 추론 및 물리 기반 실행을 분리하면서 각 모듈 간의 긴밀한 연관성을 유지하는 모듈형 특이 상황 생성 파이프라인인 CARLA-GS를 제안합니다. 실제 주행 데이터를 기반으로, 추가적인 기하학적 일관성 제약을 갖춘 편집 가능한 가우시안 장면을 재구성합니다. 다중 에이전트 LLM은 장면 수준의 추론을 수행하여 위험한 상호작용을 식별하고 의도 수준의 경로 지점을 생성하며, 저수준 동작 제어는 CARLA로 위임하여 PID 컨트롤러가 운동학적 및 동역학적 타당성을 보장합니다. 시뮬레이션된 차량 상태는 최종적으로 가우시안 장면으로 투영되어 1인칭 시점에서 렌더링됩니다. 이러한 설계는 고급 의미론적 추론, 물리적으로 실행 가능한 동작 및 사실적인 특이 상황 생성을 하나의 통합 파이프라인 내에서 가능하게 합니다. Waymo Open 데이터 세트에 대한 실험 결과는 정량적 및 정성적으로, 우리 프레임워크가 제어 가능한 특이 상황 생성 능력을 제공하며, 의미론적 의도와 물리적으로 타당한 동작에 맞춰 사실적인 공간-시간 일관성을 갖는 비디오를 생성할 수 있음을 보여줍니다.
Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations. Corner-case generation is inherently a multi-source problem spanning visual representation, scene reasoning, and vehicle trajectory generation and control. Prior knowledge- and model-based approaches typically focus on scene or trajectory components in isolation, while diffusion-based methods attempt end-to-end generation but still struggle to ensure spatiotemporal consistency and physical realism. To unify these aspects within a single framework, we propose CARLA-GS, a modular corner-case synthesis pipeline that decouples visual representation, semantic reasoning, and physics-based execution while maintaining tight cross-module coupling. Starting from real driving data, we reconstruct an editable gaussian scene with additional geometry-consistent constraints. A multi-agent LLM then performs scene-level reasoning to identify risky interactions and generate intent-level waypoint trajectories, while the low-level motion control is delegated to CARLA, where a PID controller ensures kinematic and dynamic feasibility. The simulated vehicle states are finally re-projected into the gaussian scene for ego-centric rendering. This design enables high-level semantic reasoning, low-level physically executable motion, and photorealistic corner-case generation within a unified pipeline. Experiments on the Waymo Open Dataset show, both quantitatively and qualitatively, that our framework enables controllable corner-case generation and produces photorealistic, spatiotemporally consistent videos aligned with semantic intent and physically feasible motion.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.