2606.10833v1 Jun 09, 2026 cs.AI

대규모 시각-언어 모델은 엔지니어처럼 추론할 수 있는가? 벤치마크 및 단계별 평가

Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation

Syed Mohamad Tawseeq
Syed Mohamad Tawseeq
Citations: 1
h-index: 1
S. Wasiq
S. Wasiq
Citations: 0
h-index: 0
Y. Bangde
Y. Bangde
Citations: 7
h-index: 1
Debaditya Roy
Debaditya Roy
Citations: 675
h-index: 11

시각-언어 모델(VLMs)은 일반적인 다중 모드 추론 벤치마크에서 뛰어난 성능을 보이지만, 엔지니어링 추론 능력을 수행하는 능력은 아직 충분히 탐구되지 않았습니다. 일반적인 시각 질의 응답과는 달리, 엔지니어링 문제 해결에는 기술 도면 해석, 관련 물리 법칙 선택, 그리고 물리적으로 일관된 다단계 추론 유지 등이 필요합니다. 이러한 능력은 엔지니어링 교육, 과학적 지원 및 기술 의사 결정에 사용되는 AI 시스템에서 점점 더 중요해지고 있으며, 추론 오류는 물리적으로 잘못되었지만 겉보기에는 그럴듯한 해결책을 초래할 수 있습니다. 기존 벤치마크는 주로 최종 답변만을 평가하고 중간 추론 과정에 대한 제한적인 평가만 제공합니다. 본 연구에서는 5개의 엔지니어링 분야를 포함하는 총 696개의 문제로 구성된 다중 모드 벤치마크인 EngVQA를 소개합니다. 또한, VLM이 생성한 해결책을 평가하기 위한 8단계의 자동 평가 프레임워크를 제안합니다. 이 프레임워크는 해결책의 각 단계를 독립적으로 평가하여 추론 오류에 대한 세밀한 분석을 가능하게 합니다. 본 연구에서는 개발된 평가 프레임워크를 사용하여 여러 최첨단 오픈 소스 및 상용 VLM을 벤치마킹하고, 현재 엔지니어링 추론 능력의 상당한 한계를 보여줍니다. 인간 평가 결과는 자동화된 프레임워크와 높은 상관관계를 나타냈으며(Pearson 상관계수: 0.975, 평균 절대 오차: 0.67), 이는 10점 척도로 평가되었습니다. 본 연구의 결과는 다중 모드 엔지니어링 추론 시스템의 신뢰성 있는 평가를 위해 과정 중심적인 평가가 얼마나 중요한지를 강조합니다.

Original Abstract

Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reasoning remains largely unexplored. Unlike general visual question answering, engineering problem solving requires interpreting technical diagrams, selecting governing physical principles, and maintaining physically consistent multi-step reasoning. These capabilities are increasingly important for AI systems used in engineering education, scientific assistance, and technical decision-making, where reasoning failures may produce physically invalid yet superficially plausible solutions. Existing benchmarks primarily evaluate final answers and provide limited assessment of intermediate reasoning processes. We introduce EngVQA, a multimodal benchmark for evaluating engineering reasoning across 5 engineering subjects containing 696 problems. We introduce an 8-stage automatic evaluation framework for assessing VLM-generated solutions. The framework independently evaluates each stage of the solution, enabling fine-grained analysis of reasoning failures. We benchmark multiple state-of-the-art open and closed source VLMs on our evaluation framework and demonstrate substantial limitations in current engineering reasoning capabilities. Human evaluation shows strong agreement with our automated framework, achieving a Pearson correlation of 0.975 and a mean absolute error of 0.67 on a 10-point grading scale. Our results highlight the importance of process-oriented evaluation for reliable assessment of multimodal engineering reasoning systems.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!