2607.24507v1 Jul 27, 2026 cs.LG

UNIFUSION: 자기회귀 언어 모델을 통합 역률 목표 하에 이산 확산 모델로 변환

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

Xiaoyi Jiang
Xiaoyi Jiang
Citations: 10
h-index: 1
Wei Liu
Wei Liu
Citations: 77
h-index: 4
Zuoqiang Shi
Zuoqiang Shi
Citations: 4
h-index: 1
Pipi Hu
Pipi Hu
Citations: 49
h-index: 3
Yixuan Jiang
Yixuan Jiang
Citations: 2
h-index: 1
Yi Zhu
Yi Zhu
Citations: 47
h-index: 5
Jingyuan Li
Jingyuan Li
Citations: 0
h-index: 0

기존 방법은 주로 사전 학습된 자기회귀(AR) 언어 모델을 마스킹된 확산 모델에 적용하지만, 본 연구에서는 모든 토큰이 샘플링 과정에서 편집 가능한 균일 노이즈 확산 모델에 직접적으로 적용합니다. 그러나 기존의 DLM들은 서로 다른 목표 함수와 예측 파라미터화를 사용하기 때문에, AR 체크포인트를 다양한 손상 커널에 걸쳐 적응시키는 것은 어려운 과제였습니다. 본 연구는 SEDD, MDLM/GIDD, M2S 및 Neural CTMC 간의 연관성을 확립하고, 이들의 조건부 손실을 모델 역률에 대한 단일화된 일반 Kullback-Leibler 목표 함수로 표현합니다. 또한, 클린 토큰 예측을 구체적인 점수, 사후 평균 및 탈출률/점프 파라미터화로 변환하여, 마스킹 커널과 균일 커널 간의 전환을 지원하는 공유된 x0 인터페이스를 제공합니다. 이러한 연관성을 바탕으로, 본 연구에서는 사전 학습된 GPT2 체크포인트를 균일 노이즈 확산 모델에 직접적으로 적응시키는 간단한 지속적인 사전 훈련 방법인 UNIFUSION을 제안합니다. 1억 2천만 개 및 3억 5천5백만 개의 파라미터를 가진 모델의 체계적인 평가를 통해, UNIFUSION은 샘플링 단계를 16에서 256으로 늘릴수록 생성 퍼플렉시티(GenPPL)와 유니그램 엔트로피 사이의 균형을 꾸준히 개선한다는 것을 확인했습니다. 256단계에서는 UNIFUSION-S 및 UNIFUSION-M이 각각 GenPPL/엔트로피 값인 97.783/5.2626과 71.516/5.6669를 달성했으며, 동일 규모의 다른 모델 중 어떤 것도 두 지표 모두에서 UNIFUSION보다 우수한 성능을 보이지 않았습니다. 또한, 양 규모의 모델에서 UNIFUSION은 비교된 확산 모델 중에서 가장 높은 WinoGrande, SIQA 및 BBH 정확도를 달성했습니다.

Original Abstract

Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling. However, adapting AR checkpoints across corruption kernels remains challenging because existing DLMs use different objectives and prediction parameterizations. We establish connections among SEDD, MDLM/GIDD, M2S, and Neural CTMC by expressing their conditional losses as a single generalized Kullback--Leibler objective over model reverse rates. We further derive conversions from clean-token predictions to concrete-score, posterior-mean, and exit-rate/jump parameterizations, yielding a shared \(x_0\) interface that supports switching between mask and uniform kernels. Building on these connections, we propose \ours{}, a simple continual pre-training approach for directly adapting pretrained GPT2 checkpoints to uniform-noise diffusion. Through systematic evaluation of 124M- and 355M-parameter models, we show that \ours{} steadily improves the trade-off between generative perplexity (GenPPL) and unigram entropy as the sampling budget increases from 16 to 256 steps. At 256 steps, \ours{}-S and \ours{}-M achieve GenPPL/entropy pairs of \(97.783/5.2626\) and \(71.516/5.6669\), respectively; no evaluated model at the same scale simultaneously outperforms \ours{} on both metrics. At both scales, \ours{} also achieves the highest WinoGrande, SIQA, and BBH accuracy among the compared diffusion models.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!