2608.03092v1 Aug 04, 2026 cs.LG

SMOPD: 특화 및 병합 온라인 정책 증류를 통한 다중 보상 강화 학습

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Mengyu Zhou
Mengyu Zhou
Citations: 66
h-index: 3
Xiaoxi Jiang
Xiaoxi Jiang
Citations: 142
h-index: 6
Guanjun Jiang
Guanjun Jiang
Citations: 117
h-index: 5
Yihao Liu
Yihao Liu
Citations: 327
h-index: 9
Wen Wang
Wen Wang
Citations: 0
h-index: 0
Jiahua Bao
Jiahua Bao
Citations: 0
h-index: 0
Tu Yongsiqi
Tu Yongsiqi
Citations: 0
h-index: 0
Haotian Zhou
Haotian Zhou
Citations: 114
h-index: 2
Wenkui Fan
Wenkui Fan
Citations: 0
h-index: 0

본 연구는 다중 보상 강화 학습의 모델 성능 향상을 목표로 합니다. 기존의 Group reward-Decoupled Normalization Policy Optimization (GDPO) 방법은 각 보상 차원을 개별적으로 정규화하여 집계하기 전에 보상 신호가 서로 가려지는 문제를 완화합니다. 하지만, 우리의 실험 결과는 GDPO가 여전히 다른 세분성 수준의 보상 신호를 균형 있게 처리하는 데 어려움을 겪고 있음을 보여줍니다. 특히, 특정 학습 작업에서 모델은 0.1부터 1.0까지 미세한 점수를 할당하는 밀집된 보상과 함께 0 또는 1의 이진 피드백만 제공하는 희소한 보상을 동시에 받을 수 있습니다. 이러한 경우, 희소한 보상이 충분한 최적화 신호를 제공하지 못하여 해당 능력이 효과적으로 강화되지 못할 수 있음을 발견했습니다. 따라서, 미세하게 조정된 보상에서 이미 학습된 능력을 희생하지 않고, 희소한 보상의 최적화 신호를 어떻게 강화할 수 있을까요? 이러한 한계를 극복하기 위해, 다중 보상 최적화를 위한 두 단계의 훈련 방법인 Specialize-and-Merge Online Policy Distillation (SMOPD)를 제안합니다. 1단계(특화): SMOPD는 먼저 보상 우선순위 설정을 사용하여 여러 개의 보상 특화된 '선생' 모델을 훈련하여 각 보상이 해당 신호가 최적화를 효과적으로 이끌 수 있는 조건에서 학습되도록 합니다. 2단계(병합): SMOPD는 온라인 정책 증류를 활용하여 이러한 '선생' 모델의 보상 특화 능력을 단일 '학생' 정책으로 결합하면서, 동시에 균형 잡힌 작업 수준의 최적화를 유지합니다. 제안하는 방법을 검증하기 위해, 상호 보완적인 보상(도구 사용 정확도 및 형식)과 충돌하는 보상(유용성 및 안전성)이라는 두 가지 다중 보상 환경에서 실험을 수행했습니다. 이러한 설정에 기반하여, SMOPD는 1.5B, 3B 및 7B 모델 아키텍처에서 GDPO보다 우수한 성능을 보였습니다.

Original Abstract

We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization Policy Optimization (GDPO) has mitigated the issue of reward signals masking one another during direct scalarization by normalizing each reward dimension separately before aggregation. However, our experiments show that GDPO still struggles to balance reward signals with different granularities. Specifically, in some particular training tasks, the model may receive a dense reward that assigns fine-grained scores ranging from 0.1 to 1.0, together with a sparse reward that provides only binary feedback of either 0 or 1. In such cases, we find that the sparse reward may provide an insufficient optimization signal, preventing its corresponding capability from being effectively reinforced. Therefore, how can we strengthen the optimization signal from the sparse reward without sacrificing the capability already learned from the fine-grained reward? To overcome this limitation, we propose Specialize-and-Merge Online Policy Distillation (SMOPD), a two-stage training method for multi-reward optimization. Stage1-Specialize: SMOPD first employs reward-priority configurations to train multiple reward-specialized teachers, allowing each reward to be learned under conditions where its signal can effectively drive optimization. Stage2-Merge: SMOPD then utilizes online policy distillation to combine the reward-specialized capabilities of these teachers into a single student policy, while maintaining balanced task-level optimization. To validate our method, we conduct experiments on two multi-reward settings: complementary rewards(tool-calling accuracy and format) and conflicting rewards (helpful and harmless rewards). Based on above settings, SMOPD outperforms GDPO across 1.5B, 3B and 7B backbones.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!