평균-스코어 이산 확산: 스코어 엔트로피를 위한 사후 평균 기반 노이즈 제거기
Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy
스코어 엔트로피 이산 확산(SEDD)은 제약 없는 양수 스코어 비율을 사용하여 이산 역 과정에 대한 매개변수를 설정합니다. 양수 조건은 음수가 아닌 역 점프율을 보장하지만, 베이즈 실현 가능성을 보장하지 않습니다. 즉, 노이즈가 있는 상태에서의 비율은 순수한 토큰의 사후 분포를 통해 공동으로 유도될 필요가 없습니다. 스코어-엔트로피 손실 함수는 올바른 모집단 최적값을 갖지만, 이를 벗어난 경우에는 이러한 제약을 강제하지 않습니다. 훈련된 순수 균일 SEDD 모델에서 대략 4분의 1의 전체 스코어 벡터가 좌표 박스를 위반하며, 절반 이상은 박스 내에 있지만 여전히 유효한 사후 분포와 호환되지 않습니다. 이러한 위반 사항은 유한 단계 샘플링 과정에서 음의 사전 정규화 가중치를 발생시킬 수 있습니다. 원본 스코어를 브리지 폴리토프(bridge polytope)로 투영하면 관찰된 모든 음수 가중치가 제거되고, 동일한 샘플러를 사용하면서 외부 생성 PPL이 203.6에서 175.1로 개선됩니다. 우리는 '평균-스코어(mean-to-score)' (M2S) 방법을 제안합니다. M2S는 순수한 토큰의 사후 평균을 예측하고, 정확한 커널에 의존하는 선형 매핑을 통해 이를 스코어로 변환합니다. 이 방법은 특정 조건을 만족하는 알려진 좌표별 연속 시간 마르코프 체인(CTMC)에 적용 가능합니다. 균일 노이즈 환경에서는 확률 심플렉스를 브리지 폴리토프로 매핑하고, 흡수 마스크 노이즈 환경에서는 동일한 목적 함수가 MD4와 정확히 일치하도록 설계되었습니다. 제어된 28.4M 파라미터 CIFAR-10 비교 실험에서 M2S는 테스트 BPD를 3.173에서 3.129로, FID-50k 값을 $ ext{CifarSEDDFID}$에서 $ ext{CifarMtwoSFID}$로 감소시켰습니다. 약 262B 개의 OpenWebText 토큰 데이터를 사용하여 훈련된 170M 파라미터 M2S 모델은 평가된 순수 균일 SEDD, GIDD 및 신경 CTMC 모델보다 모든 테스트 샘플링 예산에서 우수한 성능을 보였으며, 특히 128 단계 샘플링 시 생성 PPL이 183.6인 최적의 순수 균일 기준 모델보다 143.3으로 더 낮은 값을 기록했습니다.
Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees nonnegative reverse jump rates, it does not ensure Bayes realizability: ratios at a noisy state need not be jointly induced by any clean-token posterior under the forward kernel. The score-entropy loss has the correct population optimum but does not enforce this constraint away from it. In a trained pure-uniform SEDD checkpoint, roughly one quarter of complete score vectors violate the coordinate box, while more than half lie inside it yet remain materially incompatible with any valid posterior. Such violations can produce negative pre-normalization weights in finite-step sampling. Projecting raw scores onto the bridge polytope removes all observed negative weights and improves external generative PPL from $203.6$ to $175.1$ without changing the sampler. We introduce \emph{mean-to-score} (M2S), which predicts a clean-token posterior mean and converts it to the score through an exact kernel-dependent linear map. The construction applies to any known coordinate-wise continuous-time Markov chain (CTMC) satisfying a mild support condition. For uniform corruption, it maps the probability simplex onto the bridge polytope; for absorbing-mask corruption, the resulting objective recovers MD4 exactly. In a controlled 28.4M-parameter CIFAR-10 comparison, M2S lowers test BPD from $3.173$ to $3.129$ and FID-50k from $\CifarSEDDFID$ to $\CifarMtwoSFID$. A 170M-parameter M2S model trained on about 262B OpenWebText token slots outperforms the evaluated pure-uniform SEDD, GIDD, and Neural CTMC checkpoints at every tested sampling budget, reaching generative PPL $143.3$ at 128 steps versus $183.6$ for the strongest pure-uniform baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.