2607.12829v1 Jul 14, 2026 cs.LG

마스크 기반 확산 대규모 언어 모델 가속화: 효율적인 추론 기술에 대한 개관

Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

Daehoon Gwak
Daehoon Gwak
Citations: 237
h-index: 5
Junwoo Park
Junwoo Park
KAIST
Citations: 162
h-index: 6
Jaegul Choo
Jaegul Choo
Citations: 366
h-index: 11
Minhyung Lee
Minhyung Lee
Citations: 0
h-index: 0

확산 대규모 언어 모델(dLLM)은 기존의 자동 회귀 모델보다 병렬 생성 측면에서 이론적인 이점을 제공합니다. 그러나 병렬 생성만으로는 실질적인 속도 향상을 보장할 수 없습니다. 이러한 효율성을 달성하려면 확산 방식을 고려한 캐싱 및 재사용과 같은 특수한 추론 메커니즘이 필요합니다. 결과적으로, 실제 배포를 위한 필수 조건으로 추론 효율성이 강조되면서 최근 연구에서는 알고리즘, 아키텍처 및 시스템 전반에 걸쳐 가속화 기술을 적극적으로 탐구하고 있습니다. 그러나 종단 간 지연 시간은 알고리즘, 아키텍처 및 시스템 수준의 복잡한 상호작용으로 인해 발생하는 경우가 많으며, 이러한 요소들이 기존 벤치마크에서 종종 혼합되어 있어 엄격한 비교가 어렵습니다. 본 개관에서는 dLLM에 대한 통일된 지연 시간 분해 프레임워크를 소개하여 이러한 요소를 분리하고 실제 배포 환경에서의 추론 속도에 미치는 영향을 분석합니다. 이 프레임워크를 기반으로, 알고리즘 혁신, 아키텍처 및 시스템 최적화, 그리고 추론 시간 스케일링이라는 세 가지 축을 따라 가속화 기술을 분류합니다. 마지막으로, 재현 가능한 벤치마킹 지침을 제공하고 병렬 생성의 잠재력을 최대한 실현하기 위한 해결해야 할 과제를 강조합니다.

Original Abstract

Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware caching and reuse. Consequently, as inference efficiency becomes a prerequisite for practical deployment, recent research has actively explored acceleration techniques across algorithms, architectures, and systems. However, rigorous comparisons remain difficult, as end-to-end latency stems from intricate trade-offs between algorithmic, architectural, and system-level factors that are often conflated in existing benchmarks. In this survey, we introduce a unified latency decomposition framework for dLLMs to disentangle these factors and analyze their impact on inference speed in real deployments. Guided by this framework, we categorize acceleration techniques along three axes covering algorithmic innovations, architectural and system optimizations, and inference-time scaling. Finally, we provide guidelines for reproducible benchmarking and highlight open challenges for realizing the full potential of parallel generation.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!