2608.00068v1 Jul 29, 2026 cs.CV

SafeBuild-Bench: 그래프 기반 데이터 마이닝을 활용한 시간적 안정성을 고려한 건설 안전 성능 평가 기준

SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

Hui Xiong
Hui Xiong
Citations: 52
h-index: 3
Qianyi Cai
Qianyi Cai
Citations: 26
h-index: 3
Bingzhuo Zhong
Bingzhuo Zhong
Citations: 163
h-index: 8
Huizai Yao
Huizai Yao
Citations: 28
h-index: 3
Yijie Xu
Yijie Xu
Citations: 109
h-index: 7
Zi-long Wang
Zi-long Wang
Citations: 0
h-index: 0
Shuai Jiang
Shuai Jiang
Citations: 0
h-index: 0
Yi Cui
Yi Cui
Citations: 4
h-index: 2

건설 안전 모델은 작업자가 난간 없이 비계 가장자리에 서 있는 경우와 같이, 일반적인 이미지 데이터 세트에서만 인식할 수 없는 실제 발생 가능한 위험 상황들을 처리해야 합니다. 하지만 실제 현장 기록들은 중복되고, 특정 현상에 편향되어 있으며, 다양한 장소와 시기에 걸쳐 수집됩니다. 본 논문에서는 시간적 변화와 현장 조건을 반영한 현실적인 건설 안전 평가를 위한 멀티모달 대규모 언어 모델(MLLM)을 평가할 수 있는 메타데이터 기반 벤치마크인 SafeBuild-Bench를 소개합니다. 이 벤치마크는 10만 건 이상의 산업용 이미지-텍스트 데이터를 기반으로 구축되었으며, 3,000장 이상의 전문가 검증된 이미지를 포함하여 객관식 위험 식별 및 자유 형식의 위험 설명 문제로 구성되어 있습니다. 전문가 검증 과정을 효율적으로 만들기 위해, 프록시 모델의 혼동 신호와 그래프 기반 다양성을 결합하여 중복 데이터 스트림에서 유용한 데이터를 선별하는 GEMS라는 그래프 기반 멀티모달 선택 파이프라인을 개발했습니다. 공개된 명령어 튜닝 데이터 세트에서 GEMS를 통해 선택된 데이터 집합은 제한적인 데이터 규모에서도 안정성 향상에 기여합니다. SafeBuild-Bench를 사용하여 현재의 MLLM 모델들을 평가한 결과, 건설 안전 분야에 대한 신뢰할 수 있는 이해도를 갖추기에는 아직 부족하며, 최고 성능을 보이는 모델조차도 약 60% 수준의 정확도를 보였습니다. 본 논문에서는 개발된 벤치마크, 평가 스크립트 및 GEMS 코드를 https://github.com/safebuild/gems 에서 공개합니다.

Original Abstract

Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only recognize common objects in curated images. Yet real inspection archives are redundant, long-tailed, and collected across changing sites and months. We introduce SafeBuild-Bench, a metadata-driven benchmark for evaluating multimodal large language models on construction safety under realistic temporal and site variation. It is mined from 100K+ industrial image-text records and contains 3,314 task instances from over 3,000 expert-verified images, covering multiple-choice hazard identification and free-form hazard description. To make expert verification scalable, we develop GEMS, a graph-enhanced multimodal selection pipeline that combines a proxy-model confusion signal with graph-based diversity to identify informative candidates from redundant streams. On public instruction-tuning data, GEMS-selected subsets preserve robustness-oriented performance under small data budgets. On SafeBuild-Bench, current MLLMs remain far from reliable construction-safety understanding, with the best overall score near 60. We release the benchmark, evaluation scripts, and GEMS codebase at https://github.com/safebuild/gems.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!