UP-Fuse: 3D 파놉틱 분할을 위한 불확실성 기반 라이다-카메라 융합
UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation
라이다-카메라 융합은 희소한 라이다 스캔을 보완하기 위해 카메라 이미지를 활용함으로써 3D 파놉틱 분할 성능을 향상시키지만, 치명적인 장애 모드를 유발하기도 한다. 악조건에서 카메라 센서의 성능 저하 또는 고장은 인지 시스템의 신뢰성을 크게 훼손할 수 있다. 이 문제를 해결하기 위해, 우리는 카메라 센서 성능 저하, 캘리브레이션 틀어짐, 그리고 센서 고장 상황에서도 견고함을 유지하는 2D 거리뷰(range-view)에서의 새로운 불확실성 인지 융합 프레임워크인 UP-Fuse를 제안한다. 원본 라이다 데이터는 먼저 거리뷰로 투영되어 라이다 인코더에 의해 인코딩되며, 동시에 카메라 특징이 추출되어 동일한 공유 공간으로 투영된다. UP-Fuse의 핵심은 예측된 불확실성 맵을 사용하여 교차 모달 상호작용을 동적으로 조절하는 불확실성 기반 융합 모듈을 채택한 것이다. 이 맵들은 다양한 시각적 저하 조건에서 표현의 발산을 정량화하여 학습되며, 신뢰할 수 있는 시각적 단서만이 융합된 표현에 영향을 미치도록 보장한다. 융합된 거리뷰 특징은 2D 투영에 내재된 공간적 모호성을 완화하고 3D 파놉틱 분할 마스크를 직접 예측하는 새로운 하이브리드 2D-3D 트랜스포머에 의해 디코딩된다. Panoptic nuScenes, SemanticKITTI, 그리고 우리가 새롭게 도입한 Panoptic Waymo 벤치마크에 대한 광범위한 실험은 심각한 시각적 손상이나 오정렬 상황에서도 강력한 성능을 유지하는 UP-Fuse의 효용성과 견고성을 입증하며, 이는 안전이 필수적인 환경에서의 로봇 인지에 매우 적합함을 보여준다.
LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces a critical failure mode. Under adverse conditions, degradation or failure of the camera sensor can significantly compromise the reliability of the perception system. To address this problem, we introduce UP-Fuse, a novel uncertainty-aware fusion framework in the 2D range-view that remains robust under camera sensor degradation, calibration drift, and sensor failure. Raw LiDAR data is first projected into the range-view and encoded by a LiDAR encoder, while camera features are simultaneously extracted and projected into the same shared space. At its core, UP-Fuse employs an uncertainty-guided fusion module that dynamically modulates cross-modal interaction using predicted uncertainty maps. These maps are learned by quantifying representational divergence under diverse visual degradations, ensuring that only reliable visual cues influence the fused representation. The fused range-view features are decoded by a novel hybrid 2D-3D transformer that mitigates spatial ambiguities inherent to the 2D projection and directly predicts 3D panoptic segmentation masks. Extensive experiments on Panoptic nuScenes, SemanticKITTI, and our introduced Panoptic Waymo benchmark demonstrate the efficacy and robustness of UP-Fuse, which maintains strong performance even under severe visual corruption or misalignment, making it well suited for robotic perception in safety-critical settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.