2604.09025v1 Apr 10, 2026 cs.CV

시각-언어 모델을 위한 기술 기반 시각적 지리 위치 정보 추출

Skill-Conditioned Visual Geolocation for Vision-Language

Chen Yang
Chen Yang
Citations: 13
h-index: 2
Yutian Jiang
Yutian Jiang
Citations: 38
h-index: 1
Chen Wu
Chen Wu
Citations: 0
h-index: 0

시각-언어 모델(VLM)은 이미지 지리 위치 정보 추출 분야에서 유망한 능력을 보여주었지만, 여전히 체계적인 지리적 추론 능력과 자율적인 자기 발전 능력에 한계가 있습니다. 기존 방법은 주로 암묵적인 파라미터 메모리에 의존하며, 이는 종종 устаревшие 지식을 활용하고 환각적인 추론을 생성합니다. 또한, 현재의 추론 과정은 '단일' 과정이며, 추론 결과에 기반한 자기 발전을 위한 피드백 루프가 부족합니다. 이러한 문제점을 해결하기 위해, 우리는 진화하는 스킬 그래프(Skill-Graph)를 기반으로 하는 학습이 필요 없는 프레임워크인 GeoSkill을 제안합니다. 먼저, GeoSkill은 인간 전문가의 경로를 분석하여 기본적인 자연어 기반 스킬로 구성된 그래프를 초기화합니다. 추론 과정에서, GeoSkill은 현재 스킬 그래프에 의해 안내되는 직접적인 추론을 수행하는 추론 모델을 사용합니다. 지속적인 성장을 위해, 자율적인 진화 메커니즘은 더 큰 모델을 사용하여 웹 규모 데이터에서 얻은 이미지-좌표 쌍에 대한 여러 번의 추론을 수행하고, 실제 환경에서 검증된 추론 결과를 활용합니다. 이 메커니즘은 성공적인 경로와 실패한 경로를 분석하여 스킬을 반복적으로 합성하고 제거함으로써 스킬 그래프를 효과적으로 확장하고, 파라미터 업데이트 없이 지리적 편향을 수정합니다. 실험 결과, GeoSkill은 GeoRC 데이터셋에서 지리 위치 정확도와 추론의 신뢰성 모두에서 뛰어난 성능을 보이며, 다양한 외부 데이터셋에 대한 우수한 일반화 성능을 유지합니다. 또한, 우리의 자율적인 진화 메커니즘은 새로운 검증 가능한 스킬의 출현을 촉진하여, 시스템이 개별 사례 연구를 넘어 실제 지리적 지식을 인식하는 능력을 크게 향상시킵니다.

Original Abstract

Vision-language models (VLMs) have shown a promising ability in image geolocation, but they still lack structured geographic reasoning and the capacity for autonomous self-evolution. Existing methods predominantly rely on implicit parametric memory, which often exploits outdated knowledge and generates hallucinated reasoning. Furthermore, current inference is a "one-off" process, lacking the feedback loops necessary for self-evolution based on reasoning outcomes. To address these issues, we propose GeoSkill, a training-free framework based on an evolving Skill-Graph. We first initialize the graph by refining human expert trajectories into atomic, natural-language skills. For execution, GeoSkill employs an inference model to perform direct reasoning guided by the current Skill-Graph. For continuous growth, an Autonomous Evolution mechanism leverages a larger model to conduct multiple reasoning rollouts on image-coordinate pairs sourced from web-scale data and verified real-world reasoning. By analyzing both successful and failed trajectories from these rollouts, the mechanism iteratively synthesizes and prunes skills, effectively expanding the Skill-Graph and correcting geographic biases without any parameter updates. Experiments demonstrate that GeoSkill achieves promising performance in both geolocation accuracy and reasoning faithfulness on GeoRC, while maintaining superior generalization across diverse external datasets. Furthermore, our autonomous evolution fosters the emergence of novel, verifiable skills, significantly enhancing the system's cognition of real-world geographic knowledge beyond isolated case studies.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!