보상 모델은 무엇을 기억하는가?
What do Reward Models Memorize?
본 논문에서는 인간 선호도 데이터셋 두 가지를 사용하여 역분석적으로 학습된 보상 모델(RM)이 무엇을 기억하는지 연구합니다. 우리는 RM이 1) 쉬운, 높은 마진을 갖는 선호도 쌍에 잘못된 방식으로 정보 저장하고, 2) 데이터셋 특정의 단축경로(예: 모델 식별자, 사용자 샘플링 전략)를 기억하며, 3) 보이지 않는 선호도 쌍에 직면했을 때 인간 선호도의 단순한 휴리스틱 상관관계(예: 길이, 준수성)를 과잉 일반화한다는 것을 보여줍니다. 전반적으로, 본 연구의 결과는 인간 선호도 데이터를 사용하여 역분석적으로 학습된 RM이 문맥 의존적인 시나리오에서 응답 품질을 판단할 수 없는 편향된 RM을 생성한다는 것을 나타냅니다.
This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.