SkMTEB: 슬로바키아 대규모 텍스트 임베딩 벤치마크 및 모델 적응
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
본 논문에서는 서부 슬라브어족의 저자원 언어인 슬로바키아를 위한 최초의 종합적인 MTEB 스타일 텍스트 임베딩 벤치마크인 SkMTEB을 소개합니다. SkMTEB은 7가지 유형의 작업에 걸쳐 총 31개의 데이터 세트를 포함하며, 기존의 다국어 슬로바키아 벤치마크보다 약 4배 깊이 있는 분석을 제공합니다. 31개의 임베딩 모델에 대한 평가 결과, 대규모의 지시형 튜닝된 다국어 모델이 가장 뛰어난 성능을 보였으며, 기존의 슬로바키아 특화 NLU(자연어 이해) 작업용으로 학습된 모델은 임베딩 작업에는 효과적으로 적용되지 않는 것으로 나타났습니다. 효율적이고 로컬에서 배포 가능한 슬로바키아 임베딩에 대한 요구 사항을 충족하기 위해, 우리는 Multilingual E5 모델에 어휘 트리밍 및 미세 조정을 적용하여 exttt{e5-sk-small} (45M 파라미터) 및 exttt{e5-sk-large} (365M 파라미터) 모델을 개발했습니다. 최대 62%의 크기 감소에도 불구하고, 당사의 오픈 소스 모델은 독점 API와 경쟁력 있는 성능을 제공하며, 의미론적 검색 및 검색 증강 생성(RAG)을 위해 로컬에서 배포될 수 있습니다. 본 논문에서는 벤치마크, 모델, 데이터 세트 및 코드를 공개적으로 제공하여 다른 저자원 언어에 대한 재현 가능한 접근 방식을 제시하고자 합니다.
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak. Our evaluation of 31 embedding models reveals that large instruction-tuned multilingual models achieve the strongest performance, while existing Slovak-specific models trained for NLU tasks transfer poorly to embedding tasks. To address the need for efficient, locally-deployable Slovak embeddings, we develop \texttt{e5-sk-small} (45M parameters) and \texttt{e5-sk-large} (365M) by applying vocabulary trimming and fine-tuning to Multilingual E5 models. Despite size reductions of up to 62\%, our open-source models achieve competitive performance with proprietary APIs while remaining locally deployable for semantic search and retrieval-augmented generation (RAG). We release the benchmark, models, datasets, and code openly, hoping our approach offers a replicable path for other under-resourced languages.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.