PreDiff-LM: 하이브리드 어텐션을 활용한 사전 학습된 이산 마스크 확산 언어 모델
PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
이산 마스크 확산 언어 모델은 양방향 생성 및 채우기 기능을 지원하지만, 사전 학습된 자기 회귀(AR) 트랜스포머를 적용하려면 원인 관계 기반 사전 훈련과 양방향 노이즈 제거 간의 조화를 이루어야 합니다. 본 연구에서는 AR 가중치 재사용 자체의 신규성보다는 어텐션 레벨에서 이 문제를 탐구합니다. PreDiff-LM은 관측된 프롬프트 내에서는 원인 관계 어텐션을 유지하면서, 마스크 처리된 대상 영역 내에서는 완전한 양방향 어텐션을 허용합니다. GPT-2 Medium, WikiText-103, 90K 단계의 동일한 설정에서, 이 하이브리드 마스크는 동일한 AR 초기화를 사용했을 때, 균일한 양방향 어텐션에 비해 언조건부 퍼플렉시티를 34.1에서 28.7로, MAUVE 점수를 0.71에서 0.78로 향상시킵니다. 또한, 어텐션 적응은 DiffuGPT 스타일의 목적 함수 적응과 함께 사용되어 26.9의 퍼플렉시티를 달성합니다. 사전 학습된 초기화는 퍼플렉시티가 50 이하가 되기까지 필요한 단계를 약 35만 단계에서 8천 단계로 줄이지만, 동일한 규모(18.9 vs 28.7)에서는 정밀하게 조정된 AR 모델이 여전히 더 강력합니다. 퍼플렉시티 외에도, PreDiff-LM은 반복성 감소, 분포 품질 향상, 네 가지 제로샷 다운스트림 작업에서의 성능 개선 및 기존 확산 기반 모델 대비 인간 선호도 증가를 보입니다. 이러한 결과는 하이브리드 어텐션이 사전 학습된 원인 관계 기반 모델을 적응시키는 데 유용한 추가적인 메커니즘임을 보여주며, 최적화된 AR 모델과의 품질 및 추론 효율성 격차를 명확히 합니다.
Discrete masked diffusion language models support bidirectional generation and infilling, but adapting pretrained autoregressive (AR) transformers requires reconciling causal pretraining with bidirectional denoising. We study this problem at the level of attention rather than claiming AR-weight reuse itself as novel. PreDiff-LM preserves causal attention within the observed prompt while allowing full bidirectional attention within the masked target. Under a matched GPT-2 Medium, WikiText-103, 90K-step setup, this hybrid mask improves unconditional perplexity from 34.1 to 28.7 and MAUVE from 0.71 to 0.78 over uniform bidirectional attention with the same AR initialization. Attention adaptation also composes with a DiffuGPT-style objective adaptation, reaching 26.9 perplexity. Pretrained initialization reduces the steps required to reach perplexity below 50 from about 350K to 8K, although a compute-matched fine-tuned AR model remains stronger at equal scale (18.9 versus 28.7). Beyond perplexity, PreDiff-LM improves repetition, distributional quality, four zero-shot downstream tasks, and human preference over prior diffusion baselines. The results position hybrid attention as a complementary mechanism for adapting pretrained causal backbones, while making explicit the remaining quality and inference-efficiency gaps to optimized AR models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.