2608.01906v1 Aug 03, 2026 cs.CV

무인항공기(UAV) 이미지로부터 재난 후 건물 피해 평가를 위한 고급 딥러닝 기술 결합의 효과 분석

Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery

Masato Taya
Masato Taya
Citations: 6
h-index: 2
Liang Cao
Liang Cao
Citations: 0
h-index: 0
Hao Niu
Hao Niu
Citations: 48
h-index: 3
H. Ung
H. Ung
Citations: 38
h-index: 3
Guillaume Habault
Guillaume Habault
Citations: 123
h-index: 7
Roberto Legaspi
Roberto Legaspi
Citations: 58
h-index: 3

재난 이후 건물의 신속하고 정확한 피해 평가는 매우 중요하지만, 여전히 어려운 과제입니다. 무인항공기(UAV) 이미지는 피해 지역에 대한 시의적절하고 고해상도의 정보를 제공하지만, 기존의 컴퓨터 비전(CV) 모델은 종종 대규모의 어노테이션 데이터가 필요하며, 지리적 영역 및 평가 정책 전반에 걸쳐 일반화 성능이 떨어지고, 특정 학습된 작업으로 제한되는 경향이 있습니다. 대규모 시각-언어 모델(LVLM)은 강력한 추론 및 일반화 능력을 통해 유망한 대안을 제시하지만, 객체 감지 및 정확한 바운딩 박스 생성과 같은 정밀하고 저수준의 인지 작업에는 부족한 면이 있습니다. 또한, 이러한 모델은 도메인 특화된 작업에 효과적으로 미세 조정하기 위해 상당량의 데이터가 필요합니다. 본 논문에서는 건물 감지와 피해 평가를 분리하는 하이브리드 프레임워크를 제안하며, CV 모델의 정밀성과 LVLM의 추론 능력을 결합합니다. 먼저 CV 모델이 건물을 감지하고 이미지에 바운딩 박스를 생성한 후, 이를 LVLM에 전달하여 피해 분류 및 문맥적 해석을 수행합니다. 저희는 이 프레임워크를 RescueNet과 FloodNet이라는 두 가지 실제 벤치마크 데이터셋으로 평가했습니다. 특히, 제안된 프레임워크 하에서 최적의 조합은 온전한 건물, 부분적으로 손상된 건물, 완전히 파괴된 건물을 정확하게 분류하며, 개별적인 기본 모델보다 최대 2.1 R^2 포인트 향상된 성능을 보였습니다. 또한, 감지 단계에서는 제한된 양의 어노테이션 데이터만 필요합니다. 단순히 전체적인 성능 향상을 보고하는 것 외에도, 저희는 실패 시나리오 및 예외 사례에 대한 상세한 분석을 제공하여 실무자에게 실질적인 통찰력을 제공하고, 향후 연구를 위한 구체적인 방향을 제시합니다. 저희의 소스 코드와 데이터는 다음 저장소에서 연구 커뮤니티가 사용할 수 있도록 공개되어 있습니다: https://github.com/ungquanghuy-kddi/VLM_GDINO.git

Original Abstract

Rapid and accurate post-disaster building damage assessment is essential, yet remains a challenging task. Unmanned Aerial Vehicle (UAV) imagery offers a timely and high-resolution view of affected areas, but existing Computer Vision (CV) models often demand large annotated datasets, generalize poorly across geographic regions and their assessment policies, and are confined to the specific tasks they were trained for. Large Vision-Language Models (LVLMs) offer a promising alternative through their strong reasoning and generalization capabilities, but fall short on precise, low-level perception tasks such as object detection and accurate bounding box generation. Furthermore, they often require a substantial amount of data for effective fine-tuning on domain-specific tasks. In this paper, we propose a hybrid framework that decouples detection from damage assessment, combining the precision of CV models with the reasoning power of LVLMs. A CV model first detects buildings and generates bounding boxes on the image that are then passed to an LVLM for damage classification and contextual interpretation. We evaluated our framework on two real-world benchmarks: RescueNet and FloodNet. In particular, the best combination under this framework accurately counts intact, partially damaged and completely destroyed buildings, surpassing isolated baselines by up to 2.1 R^2 points, while requiring only limited annotated data for the detection stage. Beyond reporting aggregate gains, we provide a detailed analysis of failure scenarios and edge cases, offering practical insights for practitioners and concrete directions for future work. Our source code and data are publicly available to the research community via the following repository: https://github.com/ungquanghuy-kddi/VLM_GDINO.git

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!