2606.30576v1 Jun 29, 2026 cs.CV

2차원 매칭을 넘어: 지오메트리 정보를 활용한 단일 단계 통합 프레임워크를 통한 다중 시점 객체 위치 추정

Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

Haojun Xu
Haojun Xu
Citations: 8
h-index: 1
Liyao Wang
Liyao Wang
Citations: 7
h-index: 1
Ruipu Wu
Ruipu Wu
Citations: 157
h-index: 3
Lei Shi
Lei Shi
Citations: 0
h-index: 0
Linjiang Huang
Linjiang Huang
Citations: 643
h-index: 14
Si Liu
Si Liu
Citations: 156
h-index: 5

다중 시점 객체 위치 추정(CVOGL)은 쿼리 뷰(예: 지상 또는 드론 이미지)에서 지오태그가 포함된 참조 이미지(예: 위성 이미지) 내의 대상 객체를 찾는 것을 목표로 합니다. 기존 방법들은 주로 2차원 외형 매칭에 의존하며, 기하학적 메타데이터, 다양한 프롬프트 및 표준 시야 이미지가 부족한 제한적인 데이터셋으로 인해 제약을 받습니다. 이러한 복합적인 문제점을 해결하기 위해, 우리는 먼저 대규모 고품질 건물 데이터셋인 exttt{dataset}을 소개합니다. 이 데이터셋은 22만 건 이상의 지상-위성 및 드론-위성 이미지 쌍으로 구성되어 있으며, 다양한 형태의 프롬프트(점, 박스, 마스크)와 카메라 포즈 정보를 제공하여 유연한 객체 지정 및 명시적인 공간 모델링을 가능하게 합니다. 또한, 순열 불변 3차원 기반 모델인 $π^3$을 활용한 새로운 단일 단계 지오메트리 인식 위치 추정 프레임워크(GAGeo)를 제안합니다. 우리 모델은 시각적 특징, 참조 프롬프트 및 학습 가능한 태스크 토큰을 원활하게 통합하여, 상속된 3차원 사전 정보를 활용하여 단일 순방향 패스에서 바운딩 박스, 세그멘테이션 마스크 및 카메라 포즈를 동시에 예측합니다. 또한, 위성 이미지를 보편적인 기준점으로 사용하는 대조 손실 함수를 도입하여 지상 및 드론 이미지의 표현을 암묵적으로 정렬함으로써, 트리플릿 학습 데이터 없이도 지상에서 드론으로의 위치 추정을 가능하게 합니다. 광범위한 실험 결과는 우리 방법이 최첨단 기술보다 훨씬 뛰어난 성능을 보이며, 새로운 장면과 다양한 다중 시점 환경에서도 탁월한 일반화 능력을 갖춘다는 것을 보여줍니다.

Original Abstract

Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged reference image (e.g., satellite). Existing approaches heavily rely on 2D appearance matching and are constrained by limited datasets lacking geometric metadata, diverse prompts, and standard field-of-view imagery. To address these intertwined challenges, we first introduce \dataset, a large-scale, high-fidelity building dataset comprising over 220,000 ground-satellite and drone-satellite pairs. It provides multi-modal prompts (points, boxes, masks) and camera poses to enable flexible target referring and explicit spatial modeling. Furthermore, we propose a novel single-stage Geometry-Aware Geo-localization framework (GAGeo), built upon the permutation-equivariant 3D foundation model $π^3$. By seamlessly integrating visual features, referring prompts, and learnable task tokens, our model adapts the inherited 3D prior to jointly predict bounding boxes, segmentation masks, and camera poses in a single forward pass. Additionally, we introduce a contrastive loss that utilizes the satellite view as a universal anchor, implicitly aligning ground and drone representations to enable zero-shot ground-to-drone localization without requiring triplet training data. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods, exhibiting exceptional generalization ability in unseen scenes and novel cross-view setups.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!