2608.09226v1 Aug 10, 2026 cs.CV

RL 기반 경로 활용 증류: 소량의 단계로 이미지 생성하기

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

Pipei Huang
Pipei Huang
Citations: 1,937
h-index: 7
Yuhan Li
Yuhan Li
Citations: 40
h-index: 3
Fangao Zeng
Fangao Zeng
Citations: 0
h-index: 0
Sicong Kang
Sicong Kang
Citations: 25
h-index: 1
Mengfei Xu
Mengfei Xu
Citations: 0
h-index: 0
Hao Zhou
Hao Zhou
Citations: 145
h-index: 4
Wei Li
Wei Li
Citations: 36
h-index: 3
Bingbing Ni
Bingbing Ni
Citations: 68
h-index: 5

효율적인 텍스트-이미지 생성을 위해서는 강화 학습(RL) 기반 보상 정합과 소량 단계 증류가 모두 필요하지만, 이러한 절차는 일반적으로 순차적으로 수행되어 학습 비용을 증가시키고 압축 과정에서 보상 효과를 잃을 위험이 있습니다. 우리는 대신 RL의 본질적인 관점을 취합니다. 기존의 확산 강화 학습은 이미 보상으로 평가된 유한 단계 경로를 생성하며, 이 경로의 중간 상태는 샘플링의 부산물이 아니라 증류 감독에 자연스러운 정보를 제공합니다. 이러한 통찰력을 바탕으로, 우리는 REST (Reward-Enhanced Scored-Trajectory Distillation)라는 단일 단계 RL-증류 공동 학습 프레임워크를 제안합니다. 이 프레임워크는 독립적인 학생 모델을 임의의 RL 교사 모델에 연결하며, 학생 모델은 교사 모델의 변화하는 경로 추적 결과로부터 구간별로 학습하지만, 원래 교사 모델 최적화는 변경하지 않습니다. 원치 않는 낮은 보상 행동을 모방하는 것을 방지하기 위해, 우리는 Advantage-Modulated Distillation (AMD)를 추가하여 경로에 따른 이점을 부호 가중치로 변환하고 기본 증류 손실에 적용합니다. AMD는 선호되는 경로로부터의 감독을 강화하고 학생 모델이 낮은 보상 경로에서 벗어나도록 합니다. 결과적으로, REST 프레임워크는 경량이며 쉽게 통합할 수 있으며, 추가적인 이미지 추적, 별도의 증류 데이터셋 또는 적대적 학습이 필요하지 않습니다. 합성 생성, 시각적 텍스트 렌더링 및 인간 선호도 정렬 실험에서, REST는 순수한 강화 학습 교사 모델과 동등하거나 그 이상의 성능을 보이는 소량 단계 조건부 생성 (CFG) 없는 추론 기능을 제공하며, 전체적인 추가 학습 비용은 순수 RL에 비해 25% 미만입니다. REST는 RTDMD보다 DrawBench PickScore를 0.82만큼 향상시키면서도 학습 반복 횟수는 1/5로 줄였습니다.

Original Abstract

Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We instead take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide a natural source of distillation supervision rather than a disposable byproduct of sampling. Based on this insight, we propose REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage RL-distillation co-training framework that attaches a decoupled student to an arbitrary RL teacher. The student learns segment-wise from the teacher's evolving rollout trajectories while leaving the original teacher optimization unchanged. To prevent uniform imitation from preserving undesirable low-reward behaviors, we further introduce Advantage-Modulated Distillation (AMD), which transforms rollout advantages into signed weights over a base distillation loss. AMD strengthens supervision from preferred trajectories and mildly repels the student from low-reward ones. The resulting framework is lightweight and plug-and-play, requires no extra image rollouts, no separate distillation dataset, and no adversarial training. Experiments on compositional generation, visual text rendering, and human-preference alignment show that REST enables few-step CFG-free inference that matches or surpasses its 40-step RL teacher, with an overall additional training cost below 25% over pure RL. REST improves DrawBench PickScore over RTDMD by 0.82 while requiring only one-fifth of the training iterations.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!