BavGround: 바이에른 지역 문화 이해도 및 방언 능력 평가를 위한 벤치마크
BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian
대규모 언어 모델(LLM)의 문화적 평가는 종종 고자원 표준 언어에 집중되어 있어, 지역 문화와 방언 사용자들이 제대로 반영되지 않는 경우가 많습니다. 본 연구에서는 영어, 독일어 및 바이에른어를 대상으로 바이에른 지역 문화 이해도 및 방언 능력을 평가하기 위한 벤치마크인 BavGround를 소개합니다. BavGround는 각 언어별로 8가지 문화 영역에 걸쳐 총 206개의 객관식 질문을 포함하고 있으며, 이는 총 618개의 병렬 데이터셋으로 구성되어 있습니다. 이 데이터셋은 광범위하게 접근 가능한 문화 지식과 함께 신문, 역사 자료 및 전문 문헌에서 얻은 지역 특화된 지식을 모두 포함합니다. 본 연구에서는 15개의 7B-10B 오픈 소스 명령어 튜닝 모델과 하나의 폐쇄형 모델을 평가했습니다. 전반적으로 다국어 성능이 우수한 모델들이 가장 좋은 결과를 보였지만, 바이에른 관련 질문 및 자료 기반 질문에 대한 성능은 낮아 방언 및 지역 문화 지식에 대한 지속적인 어려움을 나타냅니다. 또한, 평가 프로토콜에 따라 결과가 크게 달라지는 것을 확인했습니다. 문자 순서 기반 점수 부여, 섞인 문자 순서 기반 점수 부여, 옵션 텍스트 가능성, 생성된 답변 분석, 의미 매칭 등의 방법이 서로 다른 절대 점수와 순위를 산출하며, 특히 지역 특화 모델의 경우 이러한 차이가 더욱 두드러집니다. 마지막으로, GENBA-10B 체크포인트에 대한 탐색적 분석 결과, 추가 사전 훈련은 영역별로 불균등하게 답변 내용 가능성을 개선하는 반면, 방언 능력은 여전히 상대적으로 약한 것으로 나타났습니다. BavGround는 LLM에서 문화적 표현을 평가할 때 지역 특성을 고려하고, 프로토콜에 대한 인식을 바탕으로 평가를 수행하는 데 도움이 됩니다.
Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian regional cultural grounding and dialect competence across English, German and Bavarian. BavGround contains 206 multiple-choice source questions across eight cultural domains per language, yielding 618 multi-parallel instances, with items covering both broadly accessible cultural knowledge and source-grounded regional knowledge from journalism, historical sources, and specialist literature. We evaluate fifteen 7B-10B open-weight instruction-tuned models and one closed-model reference. Strong multilingual models perform best overall, but performance drops on Bavarian items and source-grounded questions, indicating persistent difficulty with dialectal and localized cultural knowledge. We further show that conclusions depend strongly on evaluation protocol: raw answer-letter scoring, shuffled-letter scoring, option-text likelihood, generated-answer parsing, and semantic matching can produce different absolute scores and rankings, especially for regionally adapted models. Finally, an exploratory analysis of GENBA-10B checkpoints suggests that continued pretraining improves answer-content likelihood unevenly across domains, while dialect competence remains comparatively weak. BavGround supports localized, protocol-aware evaluation of cultural representation in LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.