보상 모델의 문화적 선호도 조절 최적화
Steerable Cultural Preference Optimization of Reward Models
대규모 언어 모델(LLM) 기술은 다양한 문화권의 하위 커뮤니티에 서비스를 제공하며, 각 커뮤니티가 수용할 수 있는 방식으로 작동해야 합니다. 그러나 현재까지 LLM 정렬 연구는 특정 지역의 평가자들이 보이는 통일된 응답 선호도를 예측하는 데 주로 집중되어 왔습니다. 본 논문은 보다 글로벌한 관점을 가진 정렬 모델 개발을 목표로 하며, 이는 하위 커뮤니티의 선호도를 정확하게 반영하고 어느 특정 그룹에 대한 과도한 편향을 나타내지 않도록 설계되었습니다. 우리는 이러한 목적을 위해 보상 모델을 개발하고 있으며, 다양한 문화적 선호도를 균형 있게 통합할 수 있는 새로운 보상 모델 훈련 알고리즘(SCPO)을 제시합니다. 우리의 방법은 PRISM 및 GlobalOpinionQA 데이터 세트와 7개 국가에서 기준 모델 대비 소수 그룹의 보상 모델 성능을 최대 7포인트 향상시켰습니다. 또한, SCPO는 보상 모델의 전체 데이터 미세 조정 방식보다 훈련 데이터 효율성이 최대 280% 더 높습니다. 게다가, 우리는 하위 커뮤니티의 선호도를 별도로 평가하여 편향 분석을 수행하고, 우리의 가중치 부여 방법을 통해 과도한 편향이 완화됨을 보여줍니다. 저희 코드는 다음 주소에서 확인하실 수 있습니다: https://github.com/minsik-ai/Steerable-Cultural-Preference
It is essential for large language model (LLM) technology to serve many different cultural sub-communities in a manner that is acceptable to each community. However, research on LLM alignment has so far predominantly focused on predicting a unified response preference of annotators from certain regions. This paper aims to advance the development of alignment models with a more global outlook, that are able to accurately represent the preferences of subcommunities and do not exhibit excessive bias towards any of them. We focus on the development of reward models for this purpose and present a novel reward model training algorithm (SCPO) that can incorporate diverse cultural preferences in a balanced manner. Our method results in performance increases of the minority reward model of up to 7 points over the baseline model across two datasets, PRISM and GlobalOpinionQA, and across 7 countries. SCPO is up to 280% more training data-efficient than full-data finetuning of reward models. In addition, we perform analysis of bias by separately evaluating on the preference of subcommunities and show that excessive bias is mitigated via our weighting method. Our code is available at https://github.com/minsik-ai/Steerable-Cultural-Preference
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.