2606.20177v1 Jun 18, 2026 cs.CV

원격 감지 멀티모달 대규모 언어 모델(MLLM)의 부정 이해 능력 평가 및 향상

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs

Haochen Han
Haochen Han
Citations: 14
h-index: 3
Fangming Liu
Fangming Liu
Citations: 14
h-index: 3
Alex Jinpeng Wang
Alex Jinpeng Wang
Citations: 10
h-index: 2
Jueshuang Wang
Jueshuang Wang
Citations: 0
h-index: 0

멀티모달 대규모 언어 모델(MLLM)은 다양한 원격 감지(RS) 작업에서 뛰어난 성과를 보여주었습니다. 그러나 이들의 부정 이해 능력은 아직 충분히 연구되지 않았으며, 이는 실제 응용 분야에 제약을 가져옵니다. 예를 들어, 긴급 구조대는 대피 경로 중 침수되지 않은 경로를 정확하게 파악해야 하는데, 현재 모델들은 이러한 능력이 부족합니다. 이러한 한계를 포괄적으로 연구하기 위해, 본 논문에서는 지역 수준에서 장면 수준까지의 부정 이해 능력을 평가하는 최초의 벤치마크인 RS-Neg을 소개합니다. 구체적으로, 우리는 LLM을 사용하여 다양한 부정 질의를 생성하는 자동 데이터 생성 파이프라인을 설계하고, 검증을 위한 동적 시각적 집중 모듈을 도입했습니다. 우리의 평가는 최첨단 RS MLLM들이 여전히 부정 이해에 어려움을 겪고 있으며, 환각 현상을 보이고 상당한 성능 저하를 나타낸다는 것을 보여줍니다. 이러한 격차를 해소하기 위해, 우리는 모델 최적화 과정에서 부정의 논리적 역할을 명시적으로 통합하는 새로운 테스트 시간 학습 방법인 NeFo를 제안합니다. 주목할 만하게도, 약 5%의 라벨 없는 테스트 샘플을 사용하여 NeFo는 모델의 부정 이해 능력을 크게 향상시키고, 보이지 않는 작업에 대한 강력한 일반화 성능을 보여줍니다. 논문 게재 시 코드와 데이터를 공개할 예정입니다.

Original Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability to comprehend negation remains underexplored, limiting deployment in real-world applications where models must explicitly identify what is false or absent, e.g., emergency responders need to locate non-flooded routes for evacuation. To comprehensively study this limitation, we introduce RS-Neg, the first benchmark to evaluate negation understanding across region-level to scene-level tasks. Specifically, we design an automated data generation pipeline for RS imagery, using LLMs to synthesize diverse negation queries, and introduce a dynamic visual focus module for verification. Our evaluation reveals that advanced RS MLLMs struggle with negation, exhibiting hallucinations and substantial performance degradation. To close this gap, we propose NeFo, a novel test-time learning method that explicitly incorporates the logical role of negation into the model optimization. Remarkably, using about 5\% unlabeled test samples, NeFo significantly improves the negation understanding of models and shows strong generalization to unseen tasks. Code and data will be released upon acceptance.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!