보상 편향 대체 현상: 단일 축 기반 편향 완화가 최적화 압력을 상관된 지표로 재분배
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
보상 모델의 편향을 줄이기 위한 단일 축 기반 완화 전략(예: 길이, 아첨, 스타일 등 의존성 감소)은 실제로 편향을 제거하는 대신, 관련된 다른 지표로 최적화 압력을 이동시킬 수 있습니다. 우리는 이러한 현상을 '보상 편향 대체'라고 부릅니다. 이 실패는 평가 및 정책 훈련 과정에서 측정값과 실제 최적화 목표 사이의 간극 때문에 발생합니다. 우리는 완화 전략의 결과를 체계적으로 분류하고, 성공적인 완화, 편향 대체, 과도한 보정은 순위 정확도 및 승률을 포함한 모든 감사 지표 하에서 동일한 관찰 결과를 초래한다는 것을 증명했습니다. 심지어는 실제 보상에 대한 정보를 활용할 수 있는 경우에도 마찬가지입니다. 기존의 선호 학습 완화 연구를 검토한 결과, 성공적인 완화를 입증하는 데 필요한 증거를 제시하는 방법은 발견하지 못했습니다. 정책에서 유도된 분포를 평가에 포함하고 여러 가지 편향을 동시에 추적하면 측정값과 최적화 목표 사이의 간극을 효과적으로 줄일 수 있습니다. 우리는 이러한 분석 결과를 바탕으로 실제 완화 전략 및 벤치마크 개발을 위한 구체적인 지침을 제시합니다. 우리는 언어 모델 강화 학습(RLHF)에서 길이 패널티가 GRPO 훈련 동안 의도한 대로 응답 길이를 줄이지만, 최적화 압력을 신뢰도 보정에 다시 연결시켜 정책을 과신 상태로 만들고 사실 기반의 자유 형식 정확도를 저하시키는 '보상 편향 대체' 현상을 보여줍니다. 또한, 발표된 길이 편향 제거 기법 중 하나는 감사 분포에서 보상-길이 상관관계를 0으로 만들지만, 상위 N개 모델 선택 시 네 개의 최첨단(SOTA) 보상 모델 중 세 개에서 편향을 다시 도입한다는 것을 확인했습니다. 마지막으로, 길이와 아첨 간의 연관성은 인간-LLM 심사위원 간의 의견 불일치 하에서 방향이 반전되는 현상을 관찰했습니다.
Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than eliminate it, a failure mode we call reward bias substitution. The failure is enabled by a measurement-versus-optimization gap between audit and policy-induced distributions during mitigation evaluation and policy training. We formalize mitigation outcomes into a regime taxonomy and prove that successful mitigation, bias substitution, and overcorrection produce identical observables under any audit-distribution scoring, including ranking accuracy and win-rate, even when granted oracle access to the true reward. Across published preference-learning mitigation work, no method we survey reports the evidence needed to certify successful mitigation. Augmenting evaluation with policy-induced distributions while tracking multiple biases provably closes the gap, and we translate this into actionable prescriptions for mitigation methods and benchmarks. We demonstrate bias substitution in language model RLHF, where a length penalty during GRPO training compresses responses as intended yet redirects optimization pressure onto confidence calibration, driving the policy into overconfidence while factual free-form accuracy falls. We also show a published length-debiasing operator that zeroes reward-length correlation on the audit distribution but reintroduces bias under best-of-N selection on three of four SOTA reward models, and a length-sycophancy coupling whose direction reverses under human-LLM judge disagreement.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.