2604.21102v1 Apr 22, 2026 cs.CV

스트리트 뷰 이미지를 활용한 도시 환경 및 주택 특성 평가를 위한 다중 모드 대규모 언어 모델 활용

Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery

Ming Hu
Ming Hu
Citations: 4
h-index: 1
Kuangshi Ai
Kuangshi Ai
Citations: 53
h-index: 5
Chaoli Wang
Chaoli Wang
Citations: 1,308
h-index: 20
Siyuan Yao
Siyuan Yao
Citations: 182
h-index: 8
Siavash Ghorbany
Siavash Ghorbany
Citations: 346
h-index: 11
Arnav Cherukuthota
Arnav Cherukuthota
Citations: 0
h-index: 0
Meghan Forstchen
Meghan Forstchen
Citations: 224
h-index: 3
Alexis Korotasz
Alexis Korotasz
Citations: 133
h-index: 3
Matthew L. Sisk
Matthew L. Sisk
Citations: 90
h-index: 6

본 논문에서는 대규모 언어 모델(LLM)과 구글 스트리트 뷰(GSV) 이미지를 활용하여 미국 전역의 건물 상태를 자동으로 평가하는 새로운 프레임워크를 제시합니다. 저희는 Gemma 3 27B를 소규모의 사람이 직접 레이블링한 데이터셋으로 미세 조정하여 인간의 평균 의견 점수(MOS)와 높은 상관관계를 달성했으며, SRCC 및 PLCC 지표에서 MOS 기준점을 능가하는 성능을 보였습니다. 효율성을 높이기 위해, 저희는 지식 증류 기술을 적용하여 Gemma 3 27B의 기능을 더 작은 Gemma 3 4B 모델로 이전함으로써, 3배 빠른 속도로 유사한 성능을 달성했습니다. 또한, 지식을 CNN 기반 모델(EfficientNetV2-M)과 트랜스포머 모델(SwinV2-B)에 증류하여 유사한 성능을 유지하면서 30배의 속도 향상을 얻었습니다. 더 나아가, 인간-AI 정렬 연구를 통해 LLM이 도시 환경 및 주택 특성 목록을 평가하는 능력을 조사하고, LLM 평가 결과를 통합하여 주택 소유자가 추가 분석을 수행할 수 있도록 시각화 대시보드를 개발했습니다. 본 프레임워크는 대규모 건물 상태 평가를 위한 유연하고 효율적인 솔루션을 제공하며, 최소한의 인간 레이블링 노력으로 높은 정확도를 달성할 수 있습니다.

Original Abstract

We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV) imagery. By fine-tuning Gemma 3 27B on a modest human-labeled dataset, our approach achieves strong alignment with human mean opinion scores (MOS), outperforming even individual raters on SRCC and PLCC relative to the MOS benchmark. To enhance efficiency, we apply knowledge distillation, transferring the capabilities of Gemma 3 27B to a smaller Gemma 3 4B model that achieves comparable performance with a 3x speedup. Further, we distill the knowledge into a CNN-based model (EfficientNetV2-M) and a transformer (SwinV2-B), delivering close performance while achieving a 30x speed gain. Furthermore, we investigate LLMs' capabilities for assessing an extensive list of built environment and housing attributes through a human-AI alignment study and develop a visualization dashboard that integrates LLM assessment outcomes for downstream analysis by homeowners. Our framework offers a flexible and efficient solution for large-scale building condition assessment, enabling high accuracy with minimal human labeling effort.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!