3D 의미론적 장면 완성에 대한 지리공간 우선 가이드
Geospatial-Prior Guidance for 3D Semantic Scene Completion
온보드 이미지로부터 완전한 3D 기하 구조와 의미를 추론하는 것은 여전히 어려운 과제입니다. 왜냐하면 폐색 현상과 제한적인 시야 때문에 많은 장면 영역이 제약 조건 없이 남아 있기 때문입니다. 위성 이미지는 넓은 범위의 맥락 정보를 제공하지만, 외관 정보만으로는 제한적인 구조적 지침을 제공하며, 공간 또는 시간적 불일치로 인해 신뢰성이 떨어질 수 있습니다. 본 논문에서는 지리공간 정보를 활용하여 3D 의미론적 장면 완성 작업을 수행하는 GeoScene 프레임워크를 제안합니다. GeoScene은 위성 이미지와 구조화된 OpenStreetMap 정보를 함께 사용하여 3D 의미론적 장면 완성에 대한 '소프트 프라이어(soft priors)' 역할을 합니다. GeoScene은 온보드 관측 데이터와 지리공간 가이드로부터 상호 보완적인 픽셀 단위 신뢰도 가중치를 학습하고, 이를 이용하여 관측된 영역과 관측되지 않은 영역 모두에서 특징 개선을 제어합니다. 이러한 설계는 로컬 시각적 증거를 유지하면서 온보드 시야 범위를 벗어난 대규모 도로 및 건물 구조 정보를 활용합니다. SemanticKITTI와 SSCBench-KITTI-360 데이터셋에 대한 실험 결과, GeoScene은 지리공간 우선 가이드 환경에서 일관되게 기하학적 및 의미론적 완성 성능을 향상시켰으며, 특히 대규모 정적인 객체 및 지리공간 구조를 가진 클래스에서 가장 큰 효과를 보였습니다.
Inferring complete 3D geometry and semantics from onboard images remains challenging because occlusions and restricted fields of view leave large scene regions underconstrained. Although satellite imagery provides wide-area context, appearance cues alone offer limited structural guidance and may be unreliable because of spatial or temporal discrepancies. We present GeoScene, a geospatially guided framework that jointly uses satellite imagery and structured OpenStreetMap cues as soft priors for 3D semantic scene completion. GeoScene learns complementary voxel-wise reliability weights for onboard observations and geospatial guidance, and uses them to control feature refinement in observed and unobserved regions. This design preserves local visual evidence while exploiting large-scale road and building structure beyond onboard visibility. Experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that GeoScene consistently improves both geometric and semantic completion under the geospatial-prior-assisted setting, with the most pronounced benefits for large-scale static and geospatially structured classes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.