2602.05000v1 Feb 04, 2026 cs.LG

EntRGi: 엔트로피 기반 보상 지향 학습을 통한 확산 언어 모델 개선

EntRGi: Entropy Aware Reward Guidance for Diffusion Language Models

Atula Tejaswi
Atula Tejaswi
Citations: 83
h-index: 2
Litu Rout
Litu Rout
Space Applications Centre, Indian Space Research Organisation
Citations: 2,244
h-index: 14
C. Caramanis
C. Caramanis
Citations: 12,024
h-index: 49
Sanjay Shakkottai
Sanjay Shakkottai
Citations: 405
h-index: 7
Sujay Sanghavi
Sujay Sanghavi
Citations: 12
h-index: 1

보상 지향 학습은 연속적인 확산 모델의 추론 시간 적응에 큰 성공을 거두었습니다. 이는 다운스트림 보상 모델의 기울기를 사용하여 각 노이즈 제거 단계를 업데이트하는 방식입니다. 본 연구에서는 이산적인 확산 언어 모델에 대한 보상 지향 학습을 연구합니다. 이산적인 확산 언어 모델에서는 모델의 자연스러운 출력이 이산적인 토큰이기 때문에, 이를 미분할 수 없습니다. 기존 접근 방식은 이러한 이산적인 토큰을 연속적인 값으로 대체하거나, Straight-Through Estimator와 같은 기술을 사용합니다. 본 연구에서는 이러한 방법들의 단점을 분석합니다. 첫 번째 방법은 보상 모델이 연속적인 입력으로 훈련된 적이 없기 때문에 기울기 피드백이 저하됩니다. 두 번째 방법은 이산적인 토큰에서 계산된 기울기를 사용하여 연속적인 로짓을 업데이트하기 때문에 잘못된 최적화가 수행됩니다. 본 연구의 핵심적인 혁신은 EntRGi라는 새로운 메커니즘을 도입하여 이러한 상충 관계를 극복하는 것입니다. EntRGi는 모델의 신뢰도를 사용하여 연속적인 값을 조절함으로써, 보상 지향 학습을 크게 개선하고 동시에 보상 모델에 신뢰할 수 있는 입력을 제공합니다. 본 연구에서는 70억 개의 파라미터를 가진 확산 언어 모델에 대해 3가지 다양한 보상 모델과 3가지 멀티 스킬 벤치마크를 사용하여 실험을 진행했으며, 최첨단 방법보다 일관되게 성능이 향상되었음을 확인했습니다.

Original Abstract

Reward guidance has been applied to great success in the test-time adaptation of continuous diffusion models; it updates each denoising step using the gradients from a downstream reward model. We study reward guidance for discrete diffusion language models, where one cannot differentiate through the natural outputs of the model because they are discrete tokens. Existing approaches either replace these discrete tokens with continuous relaxations, or employ techniques like the straight-through estimator. In this work, we show the downsides of both these methods. The former degrades gradient feedback because the reward model has never been trained with continuous inputs. The latter involves incorrect optimization because the gradient evaluated at discrete tokens is used to update continuous logits. Our key innovation is to go beyond this tradeoff by introducing a novel mechanism called EntRGi: Entropy aware Reward Guidance that dynamically regulates the gradients from the reward model. By modulating the continuous relaxation using the model's confidence, our approach substantially improves reward guidance while providing reliable inputs to the reward model. We empirically validate our approach on a 7B-parameter diffusion language model across 3 diverse reward models and 3 multi-skill benchmarks, showing consistent improvements over state-of-the-art methods.

0 Citations
0 Influential
24.5 Altmetric
122.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!