자동으로 보상 모델의 편향성을 탐지하는 방법
Automatically Finding Reward Model Biases
보상 모델은 대규모 언어 모델(LLM)의 후속 훈련에서 중요한 역할을 합니다. 그러나 기존 연구에 따르면 보상 모델은 길이, 형식, 환각, 아첨 등과 같은 부차적인 또는 바람직하지 않은 특성을 보상할 수 있습니다. 본 연구에서는 자연어에서 보상 모델의 편향성을 자동으로 탐지하는 연구 문제를 소개하고 분석합니다. 우리는 LLM을 사용하여 반복적으로 편향 후보를 제안하고 개선하는 간단한 방법을 제시합니다. 우리의 방법은 알려진 편향성을 복구하고 새로운 편향성을 발견할 수 있습니다. 예를 들어, 선도적인 오픈 가중치 보상 모델인 Skywork-V2-8B는 종종 중복된 간격을 가진 응답과 환각된 내용을 가진 응답을 잘못 선호하는 경향이 있다는 것을 발견했습니다. 또한, 진화적 반복이 고정된 최적의 N개 항목 검색보다 우수하다는 것을 보여주고, 합성적으로 주입된 편향성을 사용하여 파이프라인의 재현성을 검증했습니다. 본 연구가 자동화된 해석 가능성 방법을 통해 보상 모델을 개선하는 데 대한 추가 연구에 기여하기를 바랍니다.
Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format, hallucinations, and sycophancy. In this work, we introduce and study the research problem of automatically finding reward model biases in natural language. We offer a simple approach of using an LLM to iteratively propose and refine candidate biases. Our method can recover known biases and surface novel ones: for example, we found that Skywork-V2-8B, a leading open-weight reward model, often mistakenly favors responses with redundant spacing and responses with hallucinated content. In addition, we show evidence that evolutionary iteration outperforms flat best-of-N search, and we validate the recall of our pipeline using synthetically injected biases. We hope our work contributes to further research on improving RMs through automated interpretability methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.