약한 지도 학습을 이용한 항공 이미지 기반 효율적인 학교 탐지: 사전 학습 및 미세 조정
Label-Efficient School Detection from Aerial Imagery via Weakly Supervised Pretraining and Fine-Tuning
정확한 학교 탐지는 교육 정책 수립, 인프라 계획, 그리고 서비스가 부족한 지역에 인터넷 연결을 확장하는 등 다양한 교육 관련 활동을 지원하는 데 필수적입니다. 그러나 전 세계 많은 지역에서 공식 기록이 오래되었거나 불완전하며, 심지어 존재하지 않는 경우도 있습니다. 수동 매핑 작업은 가치 있지만, 많은 시간과 노력이 필요하며, 넓은 지역에 적용하기 어렵습니다. 이러한 문제를 해결하기 위해, 우리는 항공 이미지에서 학교를 탐지하는 데 필요한 인간 어노테이션의 양을 최소화하고, 전 세계적인 매핑 노력을 지원하는 약한 지도 학습 프레임워크를 제안합니다. 저희 방법은 특히 데이터가 부족한 환경에서 효과적이며, 수동 어노테이션이 매우 제한적인 경우에도 적용 가능합니다. 저희는 희소한 위치 정보와 의미론적 분할을 활용하여 인프라 마스크를 자동으로 생성하고, 이를 통해 바운딩 박스를 생성하는 자동 라벨링 파이프라인을 도입했습니다. 이렇게 자동으로 라벨링된 이미지들을 사용하여, 먼저 학교의 특징을 학습하는 초기 학습 단계를 진행하고, 이후 소량의 수동으로 라벨링된 이미지를 사용하여 기존 모델을 미세 조정합니다. 이러한 2단계 학습 파이프라인은 제한된 데이터 환경에서 학교 인프라를 대규모로 정확하게 탐지하는 데 도움이 됩니다. 저희의 실험 결과는 데이터가 부족한 환경에서 높은 객체 탐지 성능을 보여주며, 단 50개의 수동으로 라벨링된 이미지만으로도 유망한 결과를 얻을 수 있음을 입증합니다. 이 프레임워크는 효율적이고 확장 가능한 방식으로 위성 이미지를 사용하여 전 세계의 학교를 매핑함으로써 교육 및 연결성 관련 활동을 지원합니다. 저희는 모든 모델, 학습 코드, 그리고 자동으로 라벨링된 데이터를 공개하여 향후 연구와 실제적인 활용을 촉진할 예정입니다.
Accurate school detection is essential for supporting education initiatives, including infrastructure planning and expanding internet connectivity to underserved areas. However, many regions around the world face challenges due to outdated, incomplete, or unavailable official records. Manual mapping efforts, while valuable, are labor-intensive and lack scalability across large geographic areas. To address this, we propose a weakly supervised framework for school detection from aerial imagery that minimizes the need for human annotations while supporting global mapping efforts. Our method is specifically designed for low-data regimes, where manual annotations are extremely scarce. We introduce an automatic labeling pipeline that leverages sparse location points and semantic segmentation to generate infrastructure masks from which we generate bounding boxes. Using these automatically labeled images, we train our detectors on a first training stage to learn a representation of what schools look like, then using a small set of manually labeled images, we fine-tune the previously trained models on this clean dataset. This two stage training pipeline enables large-scale and strong detection in low-data setting of school infrastructure with minimal supervision. Our results demonstrate strong object detection performance, particularly in the low-data regime, where the models achieve promising results using only 50 manually labeled images, significantly reducing the need for costly annotations. This framework supports education and connectivity initiatives worldwide by providing an efficient and extensible approach to mapping schools from space. All models, training code and auto-labeled data will be publicly released to foster future research and real-world impact.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.