2605.28032v1 May 27, 2026 cs.AI

석유공학 분야 대규모 언어 모델을 위한 벤치마크: PetroBench

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

Ting Zhang
Ting Zhang
Citations: 94
h-index: 3
Yingquan Wu
Yingquan Wu
Citations: 4
h-index: 1
Hengyu Meng
Hengyu Meng
Citations: 45
h-index: 3
Peng Zhou
Peng Zhou
Citations: 59
h-index: 3
Peng Li
Peng Li
Citations: 620
h-index: 14
Xiang Wang
Xiang Wang
Citations: 77
h-index: 2
Sen Wang
Sen Wang
Citations: 5
h-index: 2

대규모 언어 모델(LLM)은 석유 산업에서 점점 더 많이 활용되고 있으며, 이는 특정 도메인에 특화된 평가 프레임워크의 필요성을 강조합니다. 본 연구는 석유공학 분야 LLM을 위한 벤치마크를 개발하며, 데이터 전처리, 품질 필터링, 그리고 다중 모델 검증이라는 세 단계로 구성됩니다. 전문가 검토를 통해, 높은 도메인 관련성과 구별력을 갖춘 표준화된 질문 은행을 구축했습니다. 이 벤치마크는 생산, 저류층, 시추 공학 분야를 포괄하며, 객관식, 참/거짓, 용어 정의, 단답형 문제 유형으로 구성된 총 1,200개의 질문으로 이루어져 있습니다. 8개의 주요 LLM을 통합 API 환경에서 평가했습니다. 결과는 모델이 주관적인 질문에서 더 높은 성능을 보였으며, 이는 사실 지식의 차별화 능력에 대한 약점을 시사합니다. 객관식 및 참/거짓 문제에 대한 최고 정확도는 각각 65.3%와 74.3%였습니다. Gemini-3-Pro, Kimi-K2.5, 그리고 Claude-Opus-4.6-Thinking 모델이 전반적으로 가장 높은 점수인 72%-74%를 기록했습니다. 모델은 생산 공학 분야에서 가장 좋은 성능을 보였으며, 저류층 공학 분야에서 가장 낮은 성능을 보였습니다. 중국 모델은 객관식 문제에서 강점을 보인 반면, 국제 모델은 단답형 문제에서 약간 더 나은 성능을 나타냈습니다. 이 벤치마크는 석유공학 분야에서 LLM을 평가하고 배포하기 위한 재현 가능하고 실용적인 참고 자료를 제공합니다.

Original Abstract

Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study develops a benchmark for LLMs in petroleum engineering, including a three-stage process of data preprocessing, quality filtering, and multi-model validation. Using expert review, a standardized question bank with strong domain relevance and discriminative capability was constructed. The benchmark covers production, reservoir, and drilling engineering, with 1,200 questions across multiple-choice, true or false, term definition, and short-answer formats. Eight mainstream LLMs were evaluated under a unified API environment. Results show that models performed better on subjective than objective questions, indicating weaknesses in factual knowledge discrimination. The highest accuracies for multiple-choice and true or false questions were 65.3% and 74.3%, respectively. Gemini-3-Pro, Kimi-K2.5, and Claude-Opus-4.6-Thinking achieved the best overall scores of 72%-74%. Models performed best in production engineering and weakest in reservoir engineering. Chinese models showed advantages in multiple-choice questions, while international models performed slightly better in short-answer questions. The benchmark provides a reproducible and practical reference for evaluating and deploying LLMs in petroleum engineering.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!