스칼라 보상 모델을 위한 잠재적 추론 경로 학습: 엔드 투 엔드 접근 방식
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
보상 모델(RM)은 강화 학습을 통해 대규모 언어 모델을 인간의 선호도에 맞추는 데 중요한 역할을 합니다. 기존의 스칼라 RM은 효율적인 확률 기반 보상 모델링을 가능하게 하지만, 복잡하거나 분포 외(OOD) 작업에는 제대로 일반화되지 않는 피상적인 신호에 의존합니다. 반면, 생성적 RM은 광범위한 추론을 활용하여 어려운 작업에서 견고성을 향상시키지만, 자연어 기반 점수는 스칼라 RM이 제공하는 수치적 유연성과 확률적 해석 가능성이 부족합니다. 최근의 연구에서는 오프라인 다중 작업 학습을 통해 두 가지 패러다임을 결합하려는 시도가 있었지만, 이러한 병렬 최적화는 생성된 추론 경로가 실제로 다운스트림 스칼라 보상 예측에 얼마나 기여하는지를 보장하지 않습니다. 이러한 불일치를 해결하기 위해, 우리는 LatentRM이라는 보상 모델링 프레임워크를 제안합니다. LatentRM은 중간 추론 경로를 이산적인 잠재 변수로 학습하여, 다운스트림 스칼라 보상의 발생 확률을 명시적으로 최대화합니다. LatentRM은 잠재 추론 공간을 온라인 방식으로 엔드 투 엔드 최적화함으로써, 심층적인 추론 기반 평가와 정확한 점수를 긴밀하게 결합합니다. 다양한 in-distribution 및 OOD 데이터셋과 RLHF 실험 결과를 통해, LatentRM이 스칼라, 생성적, 그리고 하이브리드 RM보다 선호도 모델링 및 정책 정렬 측면에서 우수한 성능을 보이며, 개방형 대화부터 복잡한 추론까지 다양한 작업에 걸쳐 이를 입증했습니다.
Reward models (RMs) are central to aligning large language models with human preferences via reinforcement learning. Although traditional scalar RMs enable efficient and probabilistic reward modeling, they rely on superficial cues that fail to generalize to complex or out-of-distribution (OOD) tasks. Conversely, generative RMs leverage extensive reasoning to improve robustness on challenging tasks, but their natural language-based scores lack the numerical flexibility and probabilistic interpretability that scalar RMs offer. While recent approaches combine both paradigms through off-policy multi-task learning, such parallel optimization does not guarantee that generated reasoning traces actively align with or benefit downstream scalar reward prediction. To address this mismatch, we propose LatentRM, a reward modeling framework that learns intermediate reasoning traces as discrete latent variables to explicitly maximize the likelihood of downstream scalar rewards. Through on-policy optimization of the latent reasoning space end-to-end, LatentRM tightly couples deep reasoning-based evaluation with precise scoring. Extensive validations on in-distribution and OOD datasets and RLHF show that LatentRM outperforms scalar, generative, and hybrid RMs on preference modeling and policy alignment across tasks ranging from open-ended conversation to complex reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.