2606.10569v1 Jun 09, 2026 cs.CL

숨겨진 합의: 인간 피드백에서의 선호도 유효성 압축

Hidden Consensus:Preference-Validity Compression in Human Feedback

Kather-ine Lee
Kather-ine Lee
Citations: 4,919
h-index: 8
Dorcas Chia Ern Chua
Dorcas Chia Ern Chua
Citations: 0
h-index: 0
Chee Seng Chan
Chee Seng Chan
Citations: 36
h-index: 2
Jiaming Tan
Jiaming Tan
Citations: 2
h-index: 1
Zhen Xue Gue
Zhen Xue Gue
Citations: 0
h-index: 0
Norzalena Abdul Hamid
Norzalena Abdul Hamid
Citations: 0
h-index: 0
A. Azmi
A. Azmi
Citations: 69
h-index: 4
Keat Mei Yeong
Keat Mei Yeong
Citations: 0
h-index: 0
Aizat Izyani binti Mujab
Aizat Izyani binti Mujab
Citations: 0
h-index: 0
Hafsa Azam
Hafsa Azam
Citations: 27
h-index: 3
Cheemin Khoo
Cheemin Khoo
Citations: 8
h-index: 1
Hansol Lim
Hansol Lim
Citations: 12
h-index: 1

표준 RLHF(강화 학습 휴먼 피드백) 파이프라인은 종종 다양한 인간 판단을 단일한 스칼라 보상 목표로 축소합니다. 우리는 이러한 축소가 구조적으로 다원적인 사회에서 정렬성을 잘못 측정할 수 있다고 주장합니다. 여기서 불일치는 문화적, 역사적, 언어적, 지역적 또는 규범적 근거에 기반한 해석을 반영하는 것일 수 있으며, 단순히 주석 오류가 아닙니다. 이를 우리는 '선호도 유효성 압축(Preference-Validity Compression)'이라고 부르며, 이는 여러 개의 타당한 응답 옵션이 단일 최적화 목표로 붕괴되는 현상입니다. 말레이시아를 분석 대상으로 삼아, 프롬프트, 응답 및 수용성 판단을 해석 틀에 따라 연결하는 선호도 이벤트 관점에서 RLHF 스타일의 피드백 집계를 분석했습니다. 20명의 참가자로부터 얻은 321개의 선호도 이벤트와 107개의 트리오-주석 프롬프트를 분석한 결과, 79%의 프롬프트에서 단일 승자 집계 방식으로는 버려질 만한 여러 개의 다수 지지 응답이 존재하며, 최고 순위 응답 간의 명백한 우세 차이는 모든 다수 지지 옵션을 고려할 때 줄어듭니다. 참가자들은 종종 여러 개의 수용 가능한 응답을 선택하며, 폐기된 응답은 실제로 일관성 있는 지역적, 실용적 또는 문화적 맥락을 반영하는 것으로 나타났습니다. 이러한 결과는 다수 집계가 이 데이터셋에서 'argmax' 수용성을 측정하는 것이며, 진정한 다원적 정렬성을 측정하지 못한다는 것을 보여줍니다. 우리는 이를 측정 유효성 문제로 간주하며, 향후 정렬 방법은 '유효성 보존 일관성(Validity-Preserving Consistency)'을 충족해야 한다고 주장합니다. 즉, 단일 보상 목표로 축소하는 대신, 다원적이고 타당한 해석 틀에 대해 안정적으로 유지되어야 합니다.

Original Abstract

Standard RLHF pipelines often reduce heterogeneous human judgments into a single scalar reward target. We argue that this reduction can mis-measure alignment in structurally plural societies, where disagreement may reflect culturally, historically, linguistically, regionally, or normatively grounded interpretations rather than annotation noise. We call this failure Preference-Validity Compression, the collapse of multiple plural-valid response options into a single optimization target. Using Malaysia as a diagnostic setting, we analyze RLHF-style feedback aggregation through preference events linking prompts, responses, and acceptability judgments across interpretive frames. Across 321 preference events from 20 participants and 107 trio-annotated prompts, 79% of prompts contain more than one majority-supported response that single-winner aggregation would discard, and apparent dominance gaps between top responses diminish when all majority-supported options are considered. Participants frequently select multiple acceptable responses, and discarded responses demonstrably reflect coherent local, practical, or cultural frames. These findings show that majority aggregation in this corpus measures argmax acceptability rather than plural alignment. We treat this as a measurement-validity issue and argue that future alignment methods should satisfy Validity-Preserving Consistency, remaining stable across plural-valid interpretive frames rather than collapsing them into a single reward target.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!