BIM-Edit: IFC 기반 건물 정보 모델링을 위한 대규모 언어 모델 성능 평가
BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling
최근 대규모 언어 모델(LLM)은 컴퓨터 지원 설계(CAD) 분야에서 텍스트 지시를 통해 디자인 결과물을 생성하는 데 활용되고 있습니다. 하지만 엔지니어링 실무에서는 새로운 형상 생성을 넘어, 기존 장면을 이해하고 정확하게 수정하며 의미와 관계를 보존하는 것이 중요합니다. 그러나 많은 CAD 성능 평가 도구는 기존 모델을 편집하는 것보다 새로운 모델을 생성하는 데 초점을 맞추고 있으며, 주로 기하학적 정확성만 평가합니다. 본 연구에서는 Industry Foundation Classes (IFC) 형식으로 표현된 건물 정보 모델(BIM)의 자연어 기반 수정 능력을 평가하기 위한 벤치마크인 BIM-Edit을 소개합니다. BIM은 기하학과 의미 및 관계 구조를 함께 포함하고 있어 까다로운 테스트 환경을 제공합니다. BIM-Edit은 11개의 실제 건물 모델과 36개의 합성 장면에서 수행되는 324개의 수정 작업으로 구성되어 있습니다. 작업은 직접, 공간, 위상 세 가지 지시 유형으로 표현되며, 명시적인 수정과 장면 기반 수정을 모두 포함합니다. 우리는 출력 결과를 기하학적 정확성, 의미적 타당성 및 위상적 일관성의 세 가지 측면에서 평가했습니다. 평가된 LLM 중에서 가장 성능이 좋은 모델도 세 가지 측정 기준에 대한 평균 점수가 49.5%에 불과하며, 어떤 모델도 3.4% 이상의 작업만 완전히 해결하지 못했습니다. 이러한 결과는 현재 LLM의 기능과 구조화된 엔지니어링 설계 워크플로우의 요구 사항 간에 상당한 격차가 있음을 보여줍니다.
Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engineering practice, this requires more than creating new geometry, models must also understand existing scenes, edit them correctly, and preserve semantics and relations. However, many CAD benchmarks focus on creating new models rather than editing existing ones, and mostly evaluate geometric correctness. We introduce BIM-Edit, a benchmark for evaluating LLMs on natural-language editing of Building Information Models (BIM) represented in the Industry Foundation Classes (IFC) format. BIM provides a challenging testbed because building models encode geometry together with semantic and relational structure. BIM-Edit contains 324 editing tasks spanning 11 realistic building models and 36 synthetic scenes. Tasks are expressed using three instruction categories - direct, spatial, and topological - covering both explicit and scene-grounded edits. We evaluate outputs along three dimensions: geometric accuracy, semantic validity, and topological consistency. Across evaluated LLMs, the best-performing model achieves only 49.5% average score across the three metrics, and no model fully solves more than 3.4% of tasks. These results demonstrate a substantial gap between current LLM capabilities and the requirements of structured engineering design workflows.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.