3D-DefectBench: 정밀한 3D 생성 결함 평가를 위한 비전-언어 모델 파이프라인의 통제된 요인 실험 연구
3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects
자동화된 평가는 시간과 비용이 많이 드는 인간 검토가 필요한 생성형 3D 시스템을 확장하는 데 필수적입니다. 그러나 자동 평가 도구의 신뢰성은 기본 비전-언어 모델(VLM)뿐만 아니라, 에셋 렌더링 방식, 제공되는 시각 정보, 작업 정의 방법 및 인간 참조 라벨 구성 방식 등 전체 평가 파이프라인에 따라 달라집니다. 본 연구에서는 VLM 기반 3D 결함 탐지 파이프라인의 체계적인 분석을 위한 벤치마크 및 프레임워크인 3D-DefectBench를 소개합니다. 이 시스템은 기하학, 텍스처, 프롬프트 준수 등 9가지 정밀한 이진 결함을 포괄하는 전반적인 평가와 쌍대 비교를 제공하며, 생성기 개발 및 평가 도구 평가에 대한 실행 가능한 진단을 제공합니다. 균형 잡힌 요인 설계 방식을 사용하여 VLM, 카메라 프로토콜, 시각 입력 및 프롬프트 스키마의 네 가지 파이프라인 요인을 84가지 실험 설정으로 변조하고 약 320만 건의 결함 판단 결과를 얻었으며, 이후 더 광범위한 최첨단 모델 세트에 대한 단계별 검증을 수행했습니다. 모델 선택은 인간 라벨과의 일치성에 가장 큰 영향을 미치는 요소이지만, 나머지 요인들도 성능에 영향을 미치고 모델 선택과 상호 작용하며 최적의 구성을 변경할 수 있습니다. 평가된 설계 공간 내에서 6개의 RGB 이미지를 사용하는 간결한 프로토콜이 더 밀도가 높은 다중 이미지 설정 및 깊이나 표면 노멀을 포함하는 입력과 비교하여 유사한 성능을 보이며, 비용 효율적인 기본 옵션으로 강력합니다. 표준화된 파이프라인 하에서, 12개의 VLM 평가 도구 중 가장 우수한 것도 여전히 숙련된 인간 라벨러의 성능에 미치지 못하며, 전문가 합의 라벨을 불안정한 라벨로 대체하면 텍스처 일관성이 크게 감소합니다. 이러한 결과는 자동화된 평가 도구가 개별 모델이 아닌 전체 파이프라인으로 평가되고 다양한 인간 참조 기준에 맞춰 조정되어야 함을 보여줍니다. 사용되는 데이터셋, 프롬프트, 예측 및 Croissant 메타데이터는 Hugging Face에서 공개됩니다.
Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliability of an automated judge depends on the entire evaluation pipeline, not only the underlying vision-language model (VLM), but also how assets are rendered, what visual evidence is provided, how the task is specified, and how human reference labels are constructed. We introduce 3D-DefectBench, a benchmark and framework for systematic analysis of VLM-based 3D defect detection pipelines. It complements holistic ratings and pairwise preferences with nine fine-grained binary defects spanning geometry, texture, and prompt adherence, providing actionable diagnostics for generator development and judge evaluation. Using a balanced factorial design, we vary four pipeline factors, VLM, camera protocol, visual input, and prompt schema, across 84 inference designs and approximately 3.2 million scored defect decisions, followed by staged validation on a broader set of frontier models. Model choice is the largest determinant of agreement with human labels, but the remaining factors also affect performance, interact with model selection, and can change the best configuration. Within the evaluated design space, a compact six-view RGB protocol performs comparably to denser multi-view settings and inputs augmented with depth or surface normals, making it a strong cost-effective default. Under this standardized pipeline, the best of 12 VLM judges still lag behind trained human labelers, while texture agreement drops sharply when expert-consensus labels are replaced by noisier silver labels. These findings show that automated judges should be evaluated as complete pipelines and calibrated across human reference regimes, rather than benchmarked only as standalone models. We release labels, prompts, predictions, and Croissant metadata on Hugging Face.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.