대규모 언어 모델에서의 선호 학습을 위한 후회 최소화 프레임워크
A Regret Minimization Framework on Preference Learning in Large Language Models
검증 가능한 보상을 활용한 강화학습(RLVR)은 작업별 검증기가 제공하는 자동 정확성 신호를 기반으로 추론 중심적인 작업에서 발전을 가능하게 했습니다. 그러나 많은 실제 언어 작업에는 신뢰할 수 있는 검증기를 구축하기 어렵기 때문에, 인간 피드백을 활용한 강화학습(RLHF)에 대한 의존도가 높아지고 있습니다. 본 논문에서는 인간 피드백이 어떻게 해석되어야 하는지에 대한 면밀한 검토가 필수적이라고 주장합니다. 우리는 '후회 기반 선호 최적화(RePO)'를 제안하며, 이는 보상 극대화 대신 '후회 최소화'를 통해 RLHF를 재구성하는 방법입니다. 인간의 선호도는 종종 즉각적인, 결과에 독립적인 효용보다는 잠재적인 결과 예측 및 대안 행동과의 반사실적 비교에 의해 형성됩니다. RePO는 이러한 구조를 모델링하여, 선호를 상대적인 최적성 부족 정도에 대한 행동 기반 평가로 표현합니다. 수학적 추론 벤치마크와 인간 선호 데이터 세트에 대한 실험 결과, 일관된 성능 향상을 보였으며, 이는 RePO가 대규모 언어 모델 학습을 위한 효과적이고 인간 친화적인 접근 방식임을 시사합니다.
Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness signals. However, many realistic language tasks are difficult to equip with reliable verifiers, motivating a growing reliance on reinforcement learning from human feedback (RLHF). In this setting, we argue that a closer examination of how human feedback should be interpreted is essential. We introduce Regret-based Preference Optimization $(\textbf{RePO})$, which reframes RLHF through $\textit{regret minimization}$ rather than reward maximization. Human preferences are often shaped by $\textit{prospective}$ anticipation of outcomes and $\textit{counterfactual}$ comparisons to alternative behaviors, rather than by immediate, outcome-independent utility. $\textbf{RePO}$ captures this structure by modeling preferences as behavior-conditioned assessments of relative suboptimality. Experiments on mathematical reasoning benchmarks and human preference datasets demonstrate consistent performance gains, indicating that $\textbf{RePO}$ is an effective and human-aligned approach for training large language models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.