안정적인 RLHF를 위한 불확실성 인지 보상 모델링
Uncertainty-Aware Reward Modeling for Stable RLHF
인간 피드백 기반 강화 학습(RLHF)은 대규모 언어 모델을 인간의 선호도 데이터에 대한 보상 모델 훈련과 정책 최적화를 통해 조정합니다. 그러나 이 파이프라인은 두 가지 근본적인 문제에 직면합니다: (1) 일반적으로 결정론적 점 추정치로 작동하는 보상 모델은 예측의 신뢰성이 낮을 때 이를 알릴 수 없으며, (2) 현대적인 그룹 기반 정책 최적화는 GRPO에서처럼 이점 계산 과정에서 보상을 균일하게 취급함으로써 신뢰할 수 없는 보상 신호를 증폭시킬 수 있습니다. 정책이 점점 더 다양한 응답을 탐색함에 따라 이러한 두 가지 한계는 중요한 취약점을 야기합니다: 신뢰할 수 없는 보상 추정치가 과도한 영향력을 행사하여 심각한 보상 해킹을 유발할 수 있습니다. 우리는 퀀타일 기반 공준 예측을 통해 보상 모델에 교정된 불확실성을 부여하고, 이질적 분산 분해를 통해 GRPO의 이점을 재가중하는 불확실성 인지 보상 모델링(UARM)을 제안합니다. HelpSteer, UltraFeedback 및 PKU-SafeRLHF에 대한 실험 결과는 UARM이 표준 GRPO 및 불확실성 무시 기반 방법과 비교하여 보상 모델의 교정 수준을 크게 향상시키고, 보상 해킹을 줄이며, 전반적인 조정 품질을 향상시킨다는 것을 보여줍니다.
Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. However, this pipeline faces two fundamental challenges: (1) reward models cannot signal when their predictions are unreliable, since they usually act as deterministic point estimators; and (2) modern group-based policy optimization can amplify unreliable reward signals, as exemplified by GRPO's uniform treatment of rewards during advantage computation. As policies explore increasingly diverse responses, these two limitations create a critical vulnerability: unreliable reward estimates may be granted disproportionate influence, triggering severe reward hacking. We propose Uncertainty-Aware Reward Modeling (UARM), which equips reward models with calibrated uncertainty via quantile-based conformal prediction and reweights GRPO advantages through heteroscedastic variance decomposition. Experiments across HelpSteer, UltraFeedback, and PKU-SafeRLHF demonstrate that UARM significantly improves reward model calibration, reduces reward hacking, and enhances downstream alignment quality compared to standard GRPO and uncertainty-agnostic baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.