IslamicTurathBench: 이슬람 학문 전통(turath) 분야의 대규모 언어 모델 평가를 위한 다중 작업, 다학제 벤치마크
IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)
대규모 언어 모델(LLM)은 질문 응답, 교육 및 연구 등 다양한 분야에서 활용되고 있으며, 특히 종교 및 문화 영역에서 전문적인 자료 전통에 의존하는 답변이 필요한 경우 더욱 그렇습니다. 그러나 이슬람 학문 분야에서는 권위 있는 학문적 전통인 'turath'에 보관된 핵심 개념, 방법론 및 논쟁에 대한 고품질의 주석 데이터가 부족합니다. 본 연구에서는 대규모 언어 모델을 고전 이슬람 학문 분야에서 평가하기 위한 다중 작업, 다학제 데이터셋인 IslamicTurathBench (ISTB)를 소개합니다. 전문가들이 개발하고 검토한 ISTB는 7개의 주요 이슬람 학문 분야에 걸쳐 12세기에 걸친 35개의 공인된 자료에서 추출한 3,465개의 질문-답변 항목으로 구성되어 있습니다. 모델의 종합적인 성능을 평가하기 위해 ISTB는 '학문적 난이도(초급, 중급, 고급)'와 '작업 형식(객관식 문제, 지문 기반 이해력, 개방형 지식 질문)'이라는 두 가지 축으로 구성되었습니다. ISTB에는 전문가 패널의 종합 점수와 10개의 시스템에 대한 제로샷 기준 결과가 포함되어 있습니다. 본 데이터셋은 역사적 맥락을 고려한 학문 분야, 학문 영역, 난이도 수준 및 질문 형식에 따른 언어 모델의 성능을 재현 가능하게 평가하는 데 활용될 수 있습니다.
Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key concepts, methods, and debates preserved in the authoritative scholarly tradition, known as turath, lack high-quality annotated resources. We introduce IslamicTurathBench (ISTB), a multi-task, multi-discipline dataset for evaluating LLMs on classical Islamic scholarship. Developed and reviewed by domain experts, ISTB contains 3,465 question-answer items drawn from 35 recognised source works spanning more than 12 centuries of scholarship across seven key fields of Islamic Studies. To enable comprehensive profiling of model capabilities, ISTB is structured along two axes: scholarly demand (Beginner, Intermediate, and Advanced) and task format (multiple-choice questions, passage-based comprehension, and open-ended knowledge questions). ISTB includes aggregated scores from a scholarly human reference panel and zero-shot baselines from ten systems. The dataset supports reproducible evaluation of language-model behaviour across source works, disciplines, scholarly-demand levels, and question formats in a historically layered scholarly domain.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.