분포적 강건성을 갖는 목록 기반 선호도 최적화
Distributionally Robust Listwise Preference Optimization
기존의 언어 모델 정렬을 위한 강건성 최적화 방법은 주로 쌍별 지도 학습에 초점을 맞추고 있으며, 강건성은 데이터셋, 프롬프트 또는 선호도 쌍 수준에서 고려됩니다. 본 연구에서는 순위 레이블 불확실성을 갖는 목록 기반 선호도 최적화를 다룹니다. 특정 프롬프트와 후보 목록이 주어졌을 때, 해당 목록에 대한 관찰된 순위는 평가자 간의 일치성 부족, 유사한 항목 간의 관계, 손실된 순위 기반 피드백 또는 보상 모델 노이즈 등으로 인해 모호할 수 있습니다. 우리는 후보 목록에 조건부로 순위 레이블을 직접적으로 강건화하는 포인트별 총 변동(total-variation) 강건성 플래킷-루스(Plackett--Luce) 목적 함수를 제안합니다. 이 강건성 손실 함수는 명목적인 PL 손실과 최악의 경우에 대한 PL 수정으로 정확하게 분해될 수 있으며, 최악의 순위는 현재의 암묵적 점수를 오름차순으로 정렬하여 얻어집니다. 이를 통해 내부 최대화 문제를 $K!$의 나열 방식에서 $O(K ext{log} K)$로 줄일 수 있습니다. 이러한 처리 가능한 구조 덕분에 강력한 오프라인 및 온라인 최적화 보장을 제공합니다. 오프라인 고정 목록 설정에서는 강건성 목적 함수가 볼록하며, 투영 확률적 부분 기울기 방법은 $O(ε^{-2})$의 샘플 복잡도로 전역 $ ext{ε}$-최적해에 도달합니다. 현재 정책에 의해 생성된 후보 목록이 있는 온라인 정책 유도 설정에서는 약한 볼록성을 가지며, $ ilde O(ε^{-2})$의 Moreau-envelope 정지성을 달성합니다. 오프라인 LLM 정렬 실험 결과, 제안된 강건성 수정은 깨끗한 레이블 조건에서 성능을 크게 유지하고 노이즈 환경에서 강건성을 향상시킵니다. 온라인 정렬에서는 보상 모델에 의해 순위가 매겨진 후보 목록 확장을 더욱 안정적으로 만들고, 보상 모델 및 외부 GPT-4 평가 지표 모두를 개선합니다.
Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We instead study listwise preference optimization under ranking-label uncertainty: given a prompt and a candidate list, the observed ranking over that list may be ambiguous due to annotator inconsistency, near-ties, lossy rankwise feedback, or reward-model noise. We propose a pointwise total-variation robust Plackett--Luce objective that directly robustifies the ranking label conditional on the candidate list. The robust loss admits an exact decomposition into the nominal PL loss plus a worst-case PL correction, and the worst-case ranking is obtained by sorting current implicit scores in ascending order, reducing the inner maximization from $K!$ enumeration to $O(K\log K)$. This tractable structure yields strong offline and online optimization guarantees. In the offline fixed-list setting, the robust objective is convex and projected stochastic subgradient reaches global $ε$-suboptimality with $O(ε^{-2})$ sample complexity. In the online policy-induced setting, where candidate lists are generated by the current policy, we establish weak convexity and $\widetilde O(ε^{-2})$ Moreau-envelope stationarity. Experiments in offline LLM alignment show that the proposed robust correction largely preserves performance under clean labels and improves robustness under noise. In online alignment, it makes reward-model-ranked candidate expansion more reliable and improves both reward-model and external GPT-4 judge metrics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.