LocationAgent: 분리 전략과 매개변수적 지식 기반 증거를 활용한 이미지 위치 추정을 위한 계층적 에이전트
LocationAgent: A Hierarchical Agent for Image Geolocation via Decoupling Strategy and Evidence from Parametric Knowledge
이미지 위치 추정은 시각적 내용을 기반으로 촬영 위치를 추론하는 것을 목표로 합니다. 근본적으로 이는 '가설-검증 주기'로 구성된 추론 과정이며, 모델이 지리적 추론 능력과 지리적 사실에 대해 증거를 검증할 수 있는 능력을 모두 갖춰야 합니다. 기존 방법들은 일반적으로 지도 학습이나 궤적 기반 강화 미세 조정을 통해 위치 지식과 추론 패턴을 정적 메모리에 내재화합니다. 결과적으로 이러한 방법들은 개방형 세계(open-world) 설정이나 동적 지식이 필요한 시나리오에서 사실적 환각(hallucination)과 일반화의 한계에 취약합니다. 이러한 문제들을 해결하기 위해, 우리는 LocationAgent라는 계층적 위치 추정 에이전트를 제안합니다. 우리의 핵심 철학은 계층적 추론 논리는 모델 내에 유지하면서, 지리적 증거의 검증은 외부 도구로 이관(offloading)하는 것입니다. 계층적 추론을 구현하기 위해 우리는 역할 분리와 문맥 압축을 사용하여 다단계 추론에서의 표류(drifting) 문제를 방지하는 RER 아키텍처(Reasoner-Executor-Recorder)를 설계했습니다. 증거 검증을 위해 우리는 위치 추론을 뒷받침하는 다양한 증거를 제공하는 단서 탐색 도구 모음을 구축했습니다. 또한, 기존 데이터셋의 데이터 유출 문제와 중국 데이터 부족 문제를 해결하기 위해, 다양한 장면 입도(granularity)와 난이도를 아우르는 이미지 위치 추정 벤치마크인 CCL-Bench(China City Location Bench)를 소개합니다. 광범위한 실험을 통해 LocationAgent가 제로샷 설정에서 기존 방법들보다 최소 30% 이상 성능이 뛰어남을 입증했습니다.
Image geolocation aims to infer capture locations based on visual content. Fundamentally, this constitutes a reasoning process composed of \textit{hypothesis-verification cycles}, requiring models to possess both geospatial reasoning capabilities and the ability to verify evidence against geographic facts. Existing methods typically internalize location knowledge and reasoning patterns into static memory via supervised training or trajectory-based reinforcement fine-tuning. Consequently, these methods are prone to factual hallucinations and generalization bottlenecks in open-world settings or scenarios requiring dynamic knowledge. To address these challenges, we propose a Hierarchical Localization Agent, called LocationAgent. Our core philosophy is to retain hierarchical reasoning logic within the model while offloading the verification of geographic evidence to external tools. To implement hierarchical reasoning, we design the RER architecture (Reasoner-Executor-Recorder), which employs role separation and context compression to prevent the drifting problem in multi-step reasoning. For evidence verification, we construct a suite of clue exploration tools that provide diverse evidence to support location reasoning. Furthermore, to address data leakage and the scarcity of Chinese data in existing datasets, we introduce CCL-Bench (China City Location Bench), an image geolocation benchmark encompassing various scene granularities and difficulty levels. Extensive experiments demonstrate that LocationAgent significantly outperforms existing methods by at least 30\% in zero-shot settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.