2608.05850v1 Aug 06, 2026 cs.CL

MameLoshnLM: 이디시어를 위한 언어 모델 및 평가 기준

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

Uri Katz
Uri Katz
Bar-Ilan University
Citations: 202
h-index: 5
Reut Tsarfaty
Reut Tsarfaty
Citations: 3
h-index: 1
Omer Goldman
Omer Goldman
Bar Ilan University
Citations: 3,884
h-index: 12
Tomasz Limisiewicz
Tomasz Limisiewicz
Charles University in Prague
Citations: 3,346
h-index: 11
Noah A. Smith
Noah A. Smith
Citations: 73
h-index: 5

본 논문에서는 이디시어를 위해 특별히 설계된 최초의 오픈 소스 80억 파라미터 언어 모델인 MameLoshnLM을 소개합니다. 풍부한 문헌적 전통에도 불구하고, 이디시어는 디지털 자료가 부족하고 신뢰할 수 있는 평가 자원이 희소하여 이디시어 언어 모델링의 발전을 저해해 왔습니다. 기존의 다국어 코퍼스와 벤치마크는 종종 언어를 제대로 반영하지 못하며, 상당량의 노이즈가 많고 기계 번역된 데이터 및 오분류된 텍스트를 포함하는 경우가 많습니다. 이러한 문제를 해결하기 위해, 우리는 현대 웹 자료와 문학 작품을 결합한 고품질 이디시어 사전 학습 코퍼스인 Oytser와 번역, 언어 분석, 정보 추출, 언어 이해 등 다양한 작업을 포괄하는 다중 작업 벤치마크인 Kashes를 소개합니다. 이러한 자원을 활용하여 Llama 3.1 8B 모델을 추가로 사전 학습시켜 MameLoshnLM을 얻었습니다. 벤치마크의 모든 작업에서 MameLoshnLM은 유사한 규모의 기존 모델보다 뛰어난 성능을 보였습니다. 분석 결과, 이러한 성능 향상은 단순히 양적인 부분뿐만 아니라, 범용 다국어 모델에 비해 MameLoshnLM이 언어를 정의하는 어휘 및 형태론적 패턴을 더 잘 반영한다는 것을 보여줍니다. 이는 저자원 언어의 경우 노이즈가 많은 웹 규모의 다국어 데이터가 가진 근본적인 문제점을 시사합니다. 본 연구 결과는 이디시어 자연어 처리 분야의 기반을 마련하고, 역사적으로 풍부하지만 디지털 자료가 부족한 언어에 대한 언어 모델 개발의 실용적인 틀을 제공합니다.

Original Abstract

We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!