전문가 정렬 기반 초안 생성 방식의 분리된 대비 디코딩
Decoupled Contrastive Decoding via Expert-Aligned Drafting
대비 디코딩(CD)은 생성 품질을 향상시키지만, 아마추어 모델을 사용하는 과정 때문에 디코딩 비용이 높습니다. 추론 디코딩을 통해 CD를 가속화하는 방법에는 제안 정렬 문제가 있습니다: 대비 신호는 초안 작성기에 영향을 미쳐야 할까요, 아니면 검증에만 사용되어야 할까요? 우리는 경량화된 특징 레벨 초안 작성 환경에서 이 질문을 연구했습니다. 두 가지 통제된 진단, 즉 동등한 Cross-alpha 훈련과 근사적인 이중 초안 작성 분해는 동일한 결론을 제시합니다: 대비 인식 초안 작성이 전문가 정렬 기반 초안 작성보다 일관되게 성능이 향상되지 않습니다. 왜냐하면 일반적으로 대비 수정은 초안 작성기의 오류보다 약하고, 재구성은 이러한 오류를 증폭시킬 수 있기 때문입니다. 우리는 전문가 정렬 방식의 경량화된 제안 생성기를 사용하고, 아마추어 모델을 변경되지 않은 CD 검증에만 사용하는 분리된 대비 디코딩(DCD) 방식을 소개합니다. 표준 추론 검증은 일반적인 CD 출력 분포를 유지합니다. 주요 8B 환경에서 EAGLE3 기반 DCD는 일반적인 CD에 비해 평균적으로 1.65배에서 1.95배의 속도 향상을 달성했으며, MMLU 제안 경로 지연 시간을 아마추어 모델 결합 방식에 비해 약 5배에서 12배까지 줄였습니다.
Contrastive Decoding (CD) improves generation quality, but its amateur-model pass makes decoding expensive. Accelerating CD with speculative decoding raises a proposal-alignment question: should the contrastive signal shape the drafter, or should it remain only in verification? We study this question in the lightweight feature-level drafter regime. Two controlled diagnostics, matched Cross-alpha training and an Approximate Dual-Drafter decomposition, give the same diagnosis: contrastive-aware drafting does not consistently improve over expert-aligned drafting because the contrastive correction is usually weaker than drafter error, and reconstruction can amplify that error. We introduce Decoupled Contrastive Decoding (DCD), which drafts with an expert-aligned lightweight proposer and applies the amateur only in unchanged CD verification. Standard speculative verification preserves the vanilla-CD output distribution. Across the main 8B settings, EAGLE3-based DCD achieves average greedy speedups of 1.65 to 1.95x over vanilla CD and reduces MMLU proposal-path latency by about 5 to 12x relative to amateur-coupled proposal paths.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.