불확실성에서 결정론으로: 레이 매칭 없이 이미지 기반의 거칠기-세밀함 시각 플로어플랜 위치 추정
From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching
시각 플로어플랜 위치 추정(Visual Floorplan Localization, FLoc)은 실내 환경에서 에고 센서 이미지를 최소한의 구조적 지도와 매칭하여 위치 정보를 획득하는 유망한 기술입니다. 그러나 모달 간 정보 불균형과 반복적인 실내 구조 때문에, 시각 기반 FLoc은 다양한 가능성을 갖는 자세 분포(multimodal pose distributions)에 의해 근본적으로 어려움을 겪습니다. 즉, 시각적으로 동일한 관찰 결과가 공간적으로 분리된 서로 다른 위치로 매핑될 수 있습니다. 기존의 레이 매칭 기반 방법들은 이러한 문제를 해결하기 위해 명시적인 희소 기하학적 또는 의미론적 '레이'를 예측하지만, 이는 필연적으로 정보 손실을 야기하며 추론 과정에서 방대한 양의 전처리 및 매칭 작업을 필요로 합니다. 본 논문에서는 중간 단계의 레이 매칭 단계를 우회하고 불확실성에서 결정론으로 나아가는 거칠기-세밀함 시각 FLoc 프레임워크를 제안합니다. 거친 단계에서는 이미지에 조건화된 자세 확산 모델을 설계하여 연속적인 다중 모드 자세 분포를 파라미터화하고, 무작위로 초기화된 자세 입자를 다양한 후보 모드로 효과적으로 유도합니다. 세밀한 단계에서는 후보 중심의 플로어플랜 이미지를 기반으로 1m 미만의 제한된 자세 잔차를 예측하는 로컬 정제기를 제안하여 구조적 모호성을 크게 줄입니다. 우리의 방법은 오프라인 지도 전처리나 테스트 시점의 조회 테이블 없이도, 글로벌 다중 가설 추적과 로컬 1m 미만 정밀도 향상을 효과적으로 균형을 맞춥니다. S3D (전체) 및 ZInD 벤치마크에서 수행한 종합적인 실험 결과는 우리의 접근 방식이 최첨단 수준의 정확도와 안정성을 달성함을 보여줍니다.
Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and repetitive indoor layouts, visual FLoc is fundamentally challenged by multimodal pose distributions, where visually identical observations map to distinct, spatially separated locations. Existing ray-matching-based methods tackle this by explicitly predicting sparse geometric or semantic rays, which inherently incur information loss and demand resource-intensive preprocessing alongside exhaustive matching during inference. In this paper, we bypass the intermediate ray-matching paradigm and propose a coarse-to-fine visual FLoc framework that progresses from uncertainty to determinism. In the coarse stage, we design an image-conditioned pose diffusion model to parameterize the continuous multimodal pose distribution, effectively routing stochastically initialized pose particles toward distinct candidate modes. In the refinement stage, we propose a localized refiner that predicts bounded sub-meter pose residuals from candidate-centered floorplan crops, where structural ambiguities are largely eliminated. Our method effectively balances global multi-hypothesis tracking and local sub-meter refinement without requiring any offline map preprocessing or test-time lookup tables. Comprehensive results on the S3D (full) and ZInD benchmarks demonstrate that our approach achieves state-of-the-art accuracy and robustness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.