더 적은 데이터, 더 나은 정렬: 선호도 최적화를 위한 데이터 중심 다중 평가자 합의 방법
Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
선호도 최적화 연구는 종종 데이터를 고정하고 학습 목표를 변경하는 방식으로 진행됩니다. 본 논문에서는 작은 규모이지만 높은 신뢰도를 가진, 정책에 부합하는 응답 집합이 얼마나 효과적인 학습 신호를 제공할 수 있는지 탐구합니다. DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimization)는 대상 정책에서 후보 응답을 생성하고, 전문 평가자를 사용하여 유용성, 사실성 및 간결성을 평가하며, 프로세스 기반 비판 교정 과정을 적용하고, 높은 합의를 보이는 바람직하거나 바람직하지 않은 예시만 유지합니다. 이 절차를 통해 54,236개의 Mistral-7B 후보 응답 중 1,871개(3.45%)가 선택되었습니다. DMAPO로 학습된 KTO 모델은 MT-Bench에서 7.50의 점수를 달성하고, text-davinci-003 기준 모델 대비 길이 제어 win rate이 95.5%이며, IFEval 프롬프트 정확도가 57.3%입니다. 독립적인 쌍대 평가에서도 DMAPO가 SimPO보다 우수한 것으로 나타났습니다. GPT-4o는 129개의 보류된 프롬프트에서 23.3점, 200개의 LMSYS-Chat 외부 데이터셋 프롬프트에서 24.0점의 순수득점을 보여주었으며, Claude Opus 4.7은 보류된 데이터셋에서 24.1점의 순수득점을 기록했습니다. 평가 모델이나 기준을 변경하면 선택되는 예시가 달라지지만, 하위 작업 성능에는 큰 영향을 미치지 않습니다. 두 번째 백본 네트워크를 사용한 연구에서도 유사하게 3.41%의 수용률을 보였지만, 성능 향상은 다소 제한적이었습니다. 이러한 실험 결과를 통해, 합의 기반 필터링은 일반적인 지침에 대한 선호도 최적화를 위한 데이터 효율적인 방법임을 알 수 있으며, 이는 추가적인 평가 작업 및 평가자 판단에 의존한다는 단점이 있습니다.
Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can provide a reliable learning signal. Our method, DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimization), generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only high-consensus desirable or undesirable examples. This procedure accepts 1,871 of 54,236 Mistral-7B candidates (3.45%). KTO trained on this set reaches 7.50 on MT-Bench, 95.5% length-controlled win rate against a text-davinci-003 reference, and 57.3% IFEval prompt accuracy. Independent pairwise evaluation also favors DMAPO over SimPO: GPT-4o yields a net win rate of 23.3 points on 129 held-out prompts and 24.0 points on 200 out-of-distribution LMSYS-Chat prompts; Claude Opus 4.7 yields 24.1 points on the held-out set. Changing the evaluator model or rubric alters the selected examples but has little effect on downstream performance. A second-backbone study yields a similar 3.41% acceptance rate, although its performance gains are more modest. Across these experiments, consensus filtering offers a data-efficient route to preference optimization for general instructions, at the cost of additional curation compute and dependence on evaluator judgments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.