RealityBridge: 편집 가능한 3D 가우시안 스플래팅 기반 자율 주행 시뮬레이션과 실제 영상 간의 격차 해소
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
안전 중심의 자율 주행 시스템 개발에 필수적인 다양한 위험 상황(long-tail hazardous scenarios)은 수집 및 재현이 어렵습니다. 편집 가능한 3D 가우시안 스플래팅(3DGS) 시뮬레이션은 실제 주행 장면을 재구성하고 제어 가능한 장면 편집 기능을 제공하여 유망한 대안으로 떠오르고 있습니다. 그러나, 편집된 3DGS 렌더링 영상은 여전히 상당한 시뮬레이션-실제 격차(Sim-to-Real gap)를 가지고 있으며, 여기에는 렌더링 오류, 저하된 전경 요소, 일관성 없는 조명 및 시간적 깜박임 등이 포함됩니다. 기존의 복원 및 비디오 생성 방법은 이러한 문제를 해결하기에 충분하지 않으며, 종종 3DGS 특유의 오류를 동시에 수정하고 시각적 현실감을 높이며 시간적 일관성을 유지하는 데 어려움을 겪습니다. 이러한 격차를 해소하기 위해, 저희는 편집된 3DGS 주행 영상에 대한 구조 보존 및 자산 인지형 시뮬레이션-실제 변환 프레임워크인 RealityBridge를 제안합니다. RealityBridge는 렌더링된 비디오, 전경 마스크, 윤곽선 지도 및 의미론적 마스크와 같은 다중 모드 제어 정보를 활용하며, 경량화된 GateNet을 사용하여 백본 레이어에 대한 적응적인 조건 할당을 수행합니다. 또한, 저희는 표적 학습 데이터를 구축하고 보상 기반의 후처리 훈련과 함께 자기 회귀 방식의 장편 비디오 훈련을 도입하여 복원 품질 향상, 시간적 안정성 확보 및 환각 현상 감소를 달성했습니다. 내부 및 공개 주행 데이터셋에 대한 광범위한 실험 결과는 RealityBridge가 기존 방법보다 렌더링 오류 제거, 조명 조정 및 장기 시퀀스의 시간적 일관성 유지 측면에서 우수한 성능을 보임을 보여줍니다.
Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect at scale. Editable 3D Gaussian Splatting (3DGS) simulation offers a scalable alternative through real-scene reconstruction and controllable editing. However, edited 3DGS-rendered videos often exhibit a significant Sim-to-Real gap, manifested as rendering artifacts, degraded foreground assets, illumination mismatch, and temporal flickering. Addressing these coupled defects requires jointly restoring local appearance, harmonizing edited content, and maintaining temporal consistency, whereas existing methods typically address only a subset of these requirements. To fill this gap, we propose RealityBridge, a video restoration and harmonization framework that converts edited 3DGS renderings into realistic driving footage while preserving simulator-defined structure, edits, and dynamics. RealityBridge conditions a video foundation model on complementary modality signals, with a lightweight GateNet adaptively controlling their injection across backbone blocks. We further develop a task-oriented curation pipeline to construct training data, and design a four-stage supervised training strategy followed by reward-guided post-training. Extensive experiments demonstrate that RealityBridge outperforms existing methods in restoration and harmonization while preserving strong temporal consistency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.