2608.08800v1 Aug 09, 2026 cs.CL

LLM 사전-사전 학습의 불안정성: 항상 도움이 되는 것은 아니다. 다국어에 대한 연구

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

Ondrej Bojar
Ondrej Bojar
Citations: 295
h-index: 6
Sofiia Riazhskykh
Sofiia Riazhskykh
Citations: 0
h-index: 0
Nam Luu
Nam Luu
Citations: 8
h-index: 1

인공 언어를 사용하여 LLM을 사전 학습하는 기술(

Original Abstract

Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33% of training tokens needed to reach a certain performance. We validate this prior result for English on a larger set of natural languages across four language families, using two different tokenizers and varying model sizes. We also relate the observed gains (or losses) in token efficiency to quantified linguistic properties of the languages, such as sentence length, morphological richness, and features of dependency syntactic trees (tree depth, number of children, number of crossing dependencies). Our empirical results indicate that the reported gains depend heavily on the experiment setup and the choice of random seed, although we can confirm the trend of stable gains with 128-Dyck pretraining of small models with the Llama tokenizer for most of the examined languages. On a general note, we argue that multiple training runs should be carried out at least for a subset of experiments to avoid the community adopting unstable approaches.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!