2606.09159v1 Jun 08, 2026 cs.CL

디퓨전 언어 모델에서의 불변성 및 독립성을 고려한 통일된 에너지 기반 디코딩

Unified Energy for Invariant and Independent Decoding in Diffusion Language Models

Yatao Bian
Yatao Bian
Citations: 31
h-index: 2
Yuchen Yan
Yuchen Yan
Citations: 33
h-index: 2
Minkai Xu
Minkai Xu
Citations: 1,116
h-index: 17
Zaiquan Yang
Zaiquan Yang
Citations: 22
h-index: 2

디퓨전 언어 모델(DLM)은 전체 시퀀스를 반복적으로 노이즈 제거하면서 병렬 텍스트 생성을 가능하게 하며, 이는 자기 회귀(AR) 방식의 디코딩보다 더 유연한 장점을 제공합니다. 그러나 기존 방법들은 토큰 간의 관계를 충분히 반영하지 못하여, 특히 병렬 처리 수준이 높아질수록 AR 기반 모델에 비해 성능 격차가 발생합니다. 본 논문에서는 이러한 성능 격차에 대한 체계적인 분석을 수행하고, 세 가지 주요 요인(모델 용량, 의존성, 불변성)을 식별합니다. 이러한 문제점을 해결하기 위해, 우리는 먼저 불변성을 처리하기 위한 불변 에너지(Inv-E)와 효율적인 샘플링 기반 추정기를 제안합니다. 또한, 독립 에너지(Ind-E)를 결합하여 모든 요인을 고려하는 통일된 에너지(Uni-E)를 얻습니다. Uni-E는 샘플링 기반 파티션 추정이 필요 없이 정확하게 계산될 수 있다는 고유한 장점을 가지고 있으며, 모델에 구애받지 않아 임의 크기의 모델에도 적용 가능합니다. 또한, Uni-E는 의존성과 불변성으로 인해 발생하는 분포 변화를 수정할 수 있음을 증명했습니다. 디퓨전 언어 모델(DLM) 및 대규모 디퓨전 언어 모델(DLLM)에 대한 광범위한 실험을 통해 제안된 Uni-E의 효과성을 입증했습니다.

Original Abstract

Diffusion Language Models (DLMs) enable parallel text generation by iteratively denoising a full sequence, offering attractive flexibility compared to auto-regressive (AR) decoding. However, existing methods fail to fully capture token relationships, leading to a performance gap relative to AR baselines, especially as the degree of parallelism increases. In this paper, we give a systematic analysis of the gap, identifying three key factors: (i) model capacity, (ii) dependency, and (iii) invariance. To address these issues, we first propose an invariant energy (Inv-E) together with an effective sampling-based estimator to handle the invariance issue. By further combining with the independent energy (Ind-E), we obtain a unified energy (Uni-E), that accounts for all these factors. Uni-E enjoys a unique advantage: it can be computed exactly without sampling-based partition estimation. Besides, Uni-E is model agnostic and can therefore be scaled to models of arbitrary size. We further prove that Uni-E can correct the distribution shift caused by dependency and invariance. Extensive experiments across Diffusion Language Models (DLMs) and Diffusion Large Language Models (DLLMs) demonstrate the effectiveness of the proposed Uni-E.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!