확대 기능을 넘어: 초고해상도 원격 감지 데이터에 대한 다중 도구 시각적 추론 학습
Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing
초고해상도(UHR) 원격 감지(RS) 이미지는 도시 규모의 장면에서 상세한 지구 관측 정보를 제공하지만, 멀티모달 대규모 언어 모델(MLLM)에게 근본적인 과제를 제시합니다. 작업과 관련된 정보는 종종 희소하고, 국소적이며, 매우 큰 시각적 맥락에 걸쳐 공간적으로 분산되어 있습니다. 자연스러운 해결책은 MLLM에 능동적인 로컬 검사를 위한 확대 도구를 제공하는 것입니다. 그러나 XLRS-Bench의 초기 연구를 통해 확대 기능이 부분적으로만 효과적이라는 것을 발견했습니다. 확대 기능은 국지적으로 복구 가능한 정보가 있는 쉬운 및 중간 수준의 작업은 해결할 수 있지만, 글로벌 검색, 다지역 비교, 경로 계획 또는 분산된 증거 기반 추론을 요구하는 어려운 경우에는 성능이 정체됩니다. 이러한 결과를 바탕으로, 우리는 단일 도구인 확대 기능을 넘어 광범위한 위성 이미지로 구축된 대규모 지리공간 멀티 도구 시각적 추론 데이터셋인 GeoMTVR을 소개합니다. GeoMTVR은 13,000개의 UHR 질의응답(VQA) 샘플을 포함하며, 각 샘플에는 상호 연결된 추론 경로, 다양한 시각적 도구 호출 및 반환된 시각적 관찰 정보가 포함되어 있어 모델이 질문 분해, 도구 선택, 지역 검사, 객체 수준 기반 매핑, 보조 시각적 추론 및 교차 도구 증거 통합을 학습할 수 있도록 지원합니다. 지도 학습 미세 조정 외에도, 우리는 최적화를 중요한 도구 사용 결정에 집중시키는 도구 주의력 중심 강화 학습 알고리즘을 제안합니다. 여기에는 언제 도구를 호출해야 하는지, 어떤 도구를 선택해야 하는지, 어디에 적용해야 하는지 및 도구 출력 결과를 어떻게 해석해야 하는지가 포함됩니다. GeoMTVR에서 지도 학습과 우리의 강화 학습 알고리즘을 결합하여 UHR RS를 위한 멀티 도구 시각적 추론 MLLM인 GeoLens를 개발했습니다. 실험 결과, GeoLens는 직접적인 추론 및 단일 도구 확대 기능을 사용하는 기본 모델보다 일관되게 더 높은 정확도를 달성하고, 증거 기반 분석 능력이 뛰어나며, 효율적인 도구 사용 경로를 제공하는 것으로 나타났습니다.
Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-relevant evidence is often sparse, local, and spatially dispersed across extremely large visual contexts. A natural solution is to equip MLLMs with zoom-in tools for active local inspection. However, through a pilot study on XLRS-Bench, we find that zoom-in is only partially effective: it resolves easy and medium-level tasks with locally recoverable evidence, but saturates on hard cases requiring global search, multi-region comparison, path planning, or dispersed-evidence reasoning. Motivated by this finding, we move beyond single-tool zoom-in and introduce GeoMTVR, a large-scale Geospatial Multi-Tool Visual Reasoning dataset built from wide-area satellite imagery. GeoMTVR contains 13K UHR VQA samples with interleaved reasoning trajectories, diverse visual tool calls, and returned visual observations, enabling models to learn question decomposition, tool selection, regional inspection, object-level grounding, auxiliary visual reasoning, and cross-tool evidence integration. Beyond supervised fine-tuning, we propose a tool-attention-focused reinforcement learning algorithm that concentrates optimization on critical tool-use decisions, including when to invoke tools, which tool to select, where to apply it, and how to interpret tool outputs. By combining SFT on GeoMTVR with our RL algorithm, we develop GeoLens, a multi-tool visual reasoning MLLM for UHR RS. Experiments show that GeoLens consistently outperforms direct reasoning and single-tool zoom-in baselines, achieving stronger accuracy, better evidence grounding, and more efficient tool-use trajectories.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.