2606.25331v1 Jun 24, 2026 cs.CL

향상된 대규모 언어 확산 모델

Improved Large Language Diffusion Models

W. Zhao
W. Zhao
Citations: 2,953
h-index: 29
Shaoxuan Xu
Shaoxuan Xu
Citations: 39
h-index: 4
Chongxuan Li
Chongxuan Li
Citations: 1,814
h-index: 12
Shen Nie
Shen Nie
Citations: 2,499
h-index: 10
Jiaxin Wen
Jiaxin Wen
Citations: 932
h-index: 15
Qiyang Min
Qiyang Min
Citations: 242
h-index: 7
Zihao Huang
Zihao Huang
Citations: 242
h-index: 6
Yuxuan Song
Yuxuan Song
Citations: 2,629
h-index: 13
Yong Shan
Yong Shan
Citations: 329
h-index: 5
Yankai Lin
Yankai Lin
Citations: 1,429
h-index: 10

최신 대규모 언어 모델은 주로 자기 회귀 분해 방식과 인과적 어텐션을 사용하여 학습됩니다. 본 연구에서는 완전 양방향 어텐션으로 처음부터 학습된 80억 파라미터의 마스크 확산 언어 모델인 iLLaDA를 제시합니다. iLLaDA는 사전 학습 단계와 지도 미세 조정(SFT) 과정 전반에 걸쳐 마스크 확산 객관 함수를 유지하며, 12조 개의 토큰으로 사전 학습을 수행하고, 250억 개의 토큰으로 구성된 지시 데이터셋으로 12 에포크 동안 미세 조정을 진행합니다. 또한 효율성을 높이기 위해 가변 길이 생성을 사용하고, 객관식 평가를 위한 신뢰도 기반 점수 부여 방식을 도입했습니다. iLLaDA는 LLaDA와 비교하여 일반적인 벤치마크에서 전반적으로 성능이 향상되었으며, 예를 들어 iLLaDA-Base 모델은 BBH에서 21.6점, ARC-Challenge에서 14.9점의 성능 향상을 보였습니다. 또한 iLLaDA-Instruct 모델은 MATH에서 14.5점, HumanEval에서 16.5점의 성능 향상을 보였습니다. 자기 회귀 학습 방식을 사용하지 않았음에도 불구하고, iLLaDA는 여러 벤치마크에서 Qwen2.5 7B와 경쟁력 있는 성능을 유지합니다. 이러한 결과는 처음부터 완전 양방향 확산 방식으로 학습하는 것이 강력한 언어 모델 개발에 있어 유망한 방법임을 보여줍니다. 모델 가중치 및 코드: https://github.com/ML-GSAI/LLaDA.

Original Abstract

Modern large language models are predominantly trained with autoregressive factorization and causal attention. We present \emph{iLLaDA}, an 8B masked diffusion language model trained from scratch with fully bidirectional attention. iLLaDA keeps the masked diffusion objective throughout pre-training and supervised fine-tuning (SFT), scaling pre-training to 12T tokens and fine-tuning on a 25B-token instruction corpus for 12 epochs. We further use variable-length generation for efficiency and introduce confidence-based scoring for multiple-choice evaluation. Compared with LLaDA, iLLaDA improves broadly across general, mathematical, and code benchmarks; for example, iLLaDA-Base improves by 21.6 points on BBH and 14.9 points on ARC-Challenge, while iLLaDA-Instruct improves by 14.5 points on MATH and 16.5 points on HumanEval. Despite its non-autoregressive training, iLLaDA also remains competitive with Qwen2.5 7B on several benchmarks. These results show that fully bidirectional diffusion training from scratch is a competitive path toward strong language models. Model weights and codes: https://github.com/ML-GSAI/LLaDA.

3 Citations
0 Influential
74.5 Altmetric
13.9 Score

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!