MameLoshnLM: 이디시어를 위한 언어 모델 및 평가 기준
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
본 논문에서는 이디시어를 위해 특별히 설계된 최초의 오픈 소스 80억 파라미터 언어 모델인 MameLoshnLM을 소개합니다. 풍부한 문헌적 전통에도 불구하고, 이디시어는 디지털 자료가 부족하고 신뢰할 수 있는 평가 자원이 희소하여 이디시어 언어 모델링의 발전을 저해해 왔습니다. 기존의 다국어 코퍼스와 벤치마크는 종종 언어를 제대로 반영하지 못하며, 상당량의 노이즈가 많고 기계 번역된 데이터 및 오분류된 텍스트를 포함하는 경우가 많습니다. 이러한 문제를 해결하기 위해, 우리는 현대 웹 자료와 문학 작품을 결합한 고품질 이디시어 사전 학습 코퍼스인 Oytser와 번역, 언어 분석, 정보 추출, 언어 이해 등 다양한 작업을 포괄하는 다중 작업 벤치마크인 Kashes를 소개합니다. 이러한 자원을 활용하여 Llama 3.1 8B 모델을 추가로 사전 학습시켜 MameLoshnLM을 얻었습니다. 벤치마크의 모든 작업에서 MameLoshnLM은 유사한 규모의 기존 모델보다 뛰어난 성능을 보였습니다. 분석 결과, 이러한 성능 향상은 단순히 양적인 부분뿐만 아니라, 범용 다국어 모델에 비해 MameLoshnLM이 언어를 정의하는 어휘 및 형태론적 패턴을 더 잘 반영한다는 것을 보여줍니다. 이는 저자원 언어의 경우 노이즈가 많은 웹 규모의 다국어 데이터가 가진 근본적인 문제점을 시사합니다. 본 연구 결과는 이디시어 자연어 처리 분야의 기반을 마련하고, 역사적으로 풍부하지만 디지털 자료가 부족한 언어에 대한 언어 모델 개발의 실용적인 틀을 제공합니다.
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.