SommBench: 언어 모델의 소믈리에 전문성 평가
SommBench: Assessing Sommelier Expertise of Language Models
대규모 언어 모델의 빠른 발전으로 인해, 이들의 다국어 및 다문화 능력에 대한 체계적인 평가가 점점 더 중요해지고 있습니다. 기존의 문화 평가 벤치마크는 주로 언어적 형태로 표현될 수 있는 기본적인 문화 지식에 초점을 맞추었습니다. 본 연구에서는 소믈리에 전문성을 평가하기 위한 다국어 벤치마크인 SommBench를 제안합니다. 소믈리에 전문성은 후각 및 미각과 깊이 관련된 영역입니다. 언어 모델은 감각적 특성에 대한 정보를 텍스트 설명만으로 학습하지만, SommBench는 이러한 텍스트 기반 지식이 전문가 수준의 감각적 판단을 모방하는 데 충분한지 테스트합니다. SommBench는 세 가지 주요 작업으로 구성됩니다. 와인 이론 질문 응답(WTQA), 와인 특징 완성(WFC), 그리고 음식-와인 페어링(FWP)입니다. SommBench는 영어, 슬로바크어, 스웨덴어, 핀란드어, 독일어, 덴마크어, 이탈리아어, 그리고 스페인어 등 다양한 언어로 제공됩니다. 이를 통해 언어 모델의 와인 전문성을 언어 능력과 분리하여 평가할 수 있습니다. 벤치마크 데이터 세트는 전문 소믈리에 및 각 언어의 원어민들과 긴밀하게 협력하여 개발되었으며, 1,024개의 와인 이론 질문-응답 문제, 1,000개의 와인 특징 완성 예제, 그리고 1,000개의 음식-와인 페어링 예제가 포함되었습니다. 우리는 가장 인기 있는 언어 모델에 대한 결과를 제공하며, 여기에는 Gemini 2.5와 같은 폐쇄형 가중치 모델과 GPT-OSS 및 Qwen 3와 같은 개방형 가중치 모델이 포함됩니다. 우리의 결과는 가장 뛰어난 모델들이 와인 이론 질문 응답에서 높은 성능을 보이는 것을 보여줍니다 (폐쇄형 가중치 모델의 경우 최대 97% 정확도). 그러나 특징 완성 (최대 65% 정확도) 및 음식-와인 페어링 (MCC 범위 0에서 0.39)은 더 어려운 과제로 나타났습니다. 이러한 결과는 SommBench가 언어 모델의 소믈리에 전문성을 평가하는 데 흥미롭고 도전적인 벤치마크임을 보여줍니다. 벤치마크는 다음 주소에서 공개적으로 이용 가능합니다: https://github.com/sommify/sommbench.
With the rapid advances of large language models, it becomes increasingly important to systematically evaluate their multilingual and multicultural capabilities. Previous cultural evaluation benchmarks focus mainly on basic cultural knowledge that can be encoded in linguistic form. Here, we propose SommBench, a multilingual benchmark to assess sommelier expertise, a domain deeply grounded in the senses of smell and taste. While language models learn about sensory properties exclusively through textual descriptions, SommBench tests whether this textual grounding is sufficient to emulate expert-level sensory judgment. SommBench comprises three main tasks: Wine Theory Question Answering (WTQA), Wine Feature Completion (WFC), and Food-Wine Pairing (FWP). SommBench is available in multiple languages: English, Slovak, Swedish, Finnish, German, Danish, Italian, and Spanish. This helps separate a language model's wine expertise from its language skills. The benchmark datasets were developed in close collaboration with a professional sommelier and native speakers of the respective languages, resulting in 1,024 wine theory question-answering questions, 1,000 wine feature-completion examples, and 1,000 food-wine pairing examples. We provide results for the most popular language models, including closed-weights models such as Gemini 2.5, and open-weights models, such as GPT-OSS and Qwen 3. Our results show that the most capable models perform well on wine theory question answering (up to 97% correct with a closed-weights model), yet feature completion (peaking at 65%) and food-wine pairing show (MCC ranging between 0 and 0.39) turn out to be more challenging. These results position SommBench as an interesting and challenging benchmark for evaluating the sommelier expertise of language models. The benchmark is publicly available at https://github.com/sommify/sommbench.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.