2608.05539v1 Aug 06, 2026 cs.CV

OmniMech: 3D 재구성을 위한 통합 멀티모달 기계 부품 벤치마크

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Jingying Zeng
Jingying Zeng
Citations: 172
h-index: 8
Ziwei Dong
Ziwei Dong
Citations: 4
h-index: 1
Qi He
Qi He
Citations: 565
h-index: 11
Kai Zhang
Kai Zhang
Citations: 0
h-index: 0
Taiting Lu
Taiting Lu
Citations: 97
h-index: 5
Kaiyuan Lin
Kaiyuan Lin
Citations: 0
h-index: 0
Sisong Bei
Sisong Bei
Citations: 10
h-index: 1
Hongxing Pan
Hongxing Pan
Citations: 0
h-index: 0
Guoliang Shi
Guoliang Shi
Citations: 0
h-index: 0
Yincheng Jin
Yincheng Jin
Citations: 3
h-index: 1
Mahanth Gowda
Mahanth Gowda
Citations: 12
h-index: 2
Mingjia Wang
Mingjia Wang
Citations: 6
h-index: 1
Zhenghao Li
Zhenghao Li
Citations: 2
h-index: 1
Jiaying Lu
Jiaying Lu
Citations: 0
h-index: 0

최근의 비전-언어 모델(VLMs)은 이미지로부터 실행 가능한 CAD 프로그램을 생성할 수 있지만, 기존 방법들은 주로 거칠고 일반적인 용도의 3D 객체를 대상으로 하며, 산업용 기계 설계에 필요한 정밀한 형상과 밀리미터 단위의 공차를 고려하기 어렵습니다. 본 논문에서는 산업 제조 데이터를 기반으로 VLMs의 실행 가능한 CAD 프로그램 생성 성능을 평가하는 최초의 백만 규모 벤치마크인 OmniMech을 소개합니다. OmniMech은 완전히 치수 및 공차가 명시된 251,000개 이상의 2D 정투영 도면과 함께, 해당 원본 CAD 모델, 멀티뷰 렌더링, 메시, STEP 및 B-rep 표현, 그리고 풍부한 의미론적 주석을 포함합니다. 이 벤치마크는 다음 네 가지 과제를 포함합니다: (1) 엔지니어링 도면으로부터 파라메트릭 CAD 프로그램 합성; (2) 기하학적 및 구조적으로 일관된 재구성을 위한 다이어그램-3D 추론; (3) 치수, 기호, 특징 명칭 및 제조 제약 조건에 대한 주석 기반 추론; (4) 시각화, 측정, CAD 실행 및 검증 도구를 활용한 에이전트 기반 추론. 실험 결과, 현재의 VLMs와 CAD 특화 모델들은 여전히 실행 가능한 프로그램 합성, 정밀한 3D 재구성, 그리고 치수 및 공차의 신뢰성 있는 적용에 어려움을 겪고 있음을 보여줍니다. 본 논문에서는 벤치마크 데이터, 평가 코드, 그리고 도구 인터페이스를 공개하여 향후 연구를 지원할 예정입니다.

Original Abstract

Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the fine-grained geometry and millimeter-level tolerances required in industrial mechanical design. We introduce OmniMech, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data. OmniMech contains more than 251,000 fully dimensioned and toleranced 2D orthographic drawings, paired with native CAD models, multi-view renderings, mesh, STEP and B-rep representations, and rich semantic annotations. The benchmark includes four tasks: (1) parametric CAD program synthesis from engineering drawings; (2) diagram-to-3D reasoning for geometrically and structurally consistent reconstruction; (3) annotation-grounded reasoning over dimensions, symbols, feature callouts, and manufacturing constraints; and (4) tool-augmented agentic reasoning using visualization, measurement, CAD execution, and verification tools. Experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of dimensions and tolerances. We will release the benchmark data, evaluation code, and tool interfaces to support future research.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!