2608.02942v1 Aug 03, 2026 cs.CL

OPTD: 일관성 기반 적응형 압축을 통한 온라인 방식 전환 증류 - 소규모 단계 확산 언어 모델

OPTD: On-Policy Transition Distillation with Consistency-Guided Adaptive Compression for Few-Step Diffusion Language Models

Xiaoyi Pang
Xiaoyi Pang
Citations: 868
h-index: 14
Jian Liu
Jian Liu
Citations: 126
h-index: 4
Bohai Gu
Bohai Gu
Citations: 37
h-index: 2
Haoxi Li
Haoxi Li
Citations: 75
h-index: 5
Haoxuan Che
Haoxuan Che
Citations: 37
h-index: 3
Jingcai Guo
Jingcai Guo
Citations: 777
h-index: 18
Jiewei Zhang
Jiewei Zhang
Citations: 286
h-index: 10
Song Guo
Song Guo
Citations: 19
h-index: 2
Hualei Zhang
Hualei Zhang
Citations: 116
h-index: 4
Xiaocheng Lu
Xiaocheng Lu
Citations: 358
h-index: 9
Shuhan Guo
Shuhan Guo
Citations: 0
h-index: 0

확산 언어 모델(dLLMs)은 여러 토큰을 병렬로 예측할 수 있지만, 정확한 생성을 위해서는 여전히 많은 반복적인 노이즈 제거 단계를 필요로 합니다. 소규모 단계 증류는 여러 개의 교사 모델 단계를 하나의 학생 모델 전환으로 압축하여 디코딩 속도를 가속화합니다. 그러나 기존 방법은 오프라인 정책 경로를 기반으로 지도 정보를 구성합니다. 추론 시, 학생 모델의 초기 병렬 예측은 이후 예측의 문맥을 변경하여, 실제로 방문하는 상태가 지도된 상태와 달라지게 됩니다. 특히 단계 압축이 가장 공격적인 시점에서 이러한 불일치가 발생합니다. 온-라인 방식 증류는 이러한 불일치를 해결할 수 있는 자연스러운 방법이지만, 각 전환이 얼마나 발전해야 하는지에 대한 문제는 여전히 남아 있습니다. 교사 모델의 다음 액션에만 맞추면 압축률이 제한되고, 무분별하게 미래 액션을 병합하면 중간 의존성이 손상될 수 있습니다. 이러한 한계를 극복하기 위해, 우리는 일관성 기반 적응형 압축을 통한 온-라인 방식 전환 증류인 OPTD를 제안합니다. OPTD는 소규모 단계 학생 모델의 자체 경로에서 부분적인 상태를 샘플링하고, 고정된 질문만 입력받는 교사 모델을 사용하여 결과와 일치하는 미래 후보들을 식별하며, 현재 상태에 대한 신뢰도를 기준으로 정렬합니다. 이 방법은 교사 모델의 실행 결과를 유지하는 가장 긴 접두사를 선택합니다. '셋-병목' 목표 함수는 검증된 모든 미래 후보를 디코더의 해제 임계값으로 끌어올리고, 고정된 교사 모델 KL 앵커 규제는 나머지 활성 위치들을 정규화합니다. 제안하는 방법은 금반응 데이터를 사용하지 않고 지도 정보 구축 및 학습을 수행합니다. 네 가지 수학적 추론 및 코드 생성 벤치마크에서 OPTD는 일관되게 품질-효율 균형을 개선하며, 평가된 소규모 단계 기반 모델 중에서 가장 높은 품질 제한 AUP를 달성했습니다.

Original Abstract

Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates decoding by compressing multiple teacher steps into a single student transition. However, existing methods construct supervision on off-policy trajectories. At inference, the student's early parallel commitments alter the context of later predictions, so the states it actually visits drift away from the supervised ones--precisely when step compression is most aggressive. On-policy distillation is a natural remedy for this mismatch, but it leaves open how far each transition should advance: matching only the teacher's next action limits compression, while indiscriminately merging future actions can violate intermediate dependencies. To address this limitation, we propose OPTD, On-Policy Transition Distillation with consistency-guided adaptive compression. It samples partial states from the few-step student's own trajectories, uses a frozen, question-only teacher to identify outcome-aligned future candidates, and orders them by current-state confidence. The method then selects the longest prefix whose joint commitment preserves the teacher's rollout outcome. A set-bottleneck objective promotes every verified future candidate to the decoder's release threshold, while a frozen-teacher KL anchor regularizes all other active positions. Neither target construction nor training uses a gold response. Across four mathematical reasoning and code-generation benchmarks, OPTD consistently improves the quality--efficiency trade-off and attains the strongest overall quality-constrained AUP among the evaluated few-step baselines.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!