아첨하는 합의에서 다원적인 수정을 향하여: 인공지능 정렬이 왜 의견 불일치를 드러내야 하는가
From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement
다원적인 정렬은 일반적으로 선호도 집계 방식으로 구현되는데, 이는 다양한 인간 가치를 포괄(Overton), 특정 방향으로 유도(Steerable), 또는 비례적으로 반영(Distributional)하는 응답을 생성하는 것을 의미합니다. 우리는 집계만으로는 실제 다원적인 정렬을 구현하는 데 충분하지 않다고 주장합니다. 진정한 가치 다원주의 하에서, 현재 강화 학습 기반의 어시스턴트 모델의 실패 원인은 충분한 범위의 가치 포섭 부족이 아니라, 아첨하는 합의 경향입니다. 즉, 즉각적인 상호작용자와 동의하고, 그의 의견을 긍정하며, 갈등을 최소화하려는 학습된 경향입니다. 배포된 인공지능 시스템이 이제 건강, 시민 생활, 노동, 그리고 거버넌스와 같은 중요한 영역의 의사 결정을 중재하는 상황에서, 상호작용 단계에서의 의견 불일치의 붕괴는 단순한 기술적인 문제가 아니라, 불평등한 결과를 초래하는 구조적인 실패입니다. 우리는 그라이스의 격률에서 영감을 받은 세 가지 대화 메커니즘, 즉 관점의 한계를 명시하는 '범위 설정(scoping)', 가치 충돌을 숨기기보다 드러내는 '신호(signalling)', 그리고 원칙에 기반하여 자신의 입장을 수정하는 '수정(repair)'을 중심으로 다원적인 정렬을 재구성합니다. 우리는 원칙에 기반한 수정과 굴복을 구별하는 지표인 '다원적 수정 점수(Pluralistic Repair Score, PRS)'를 공식화하고, Claude Sonnet 4.5 (N=198) 및 GPT-4o (N=100) 두 가지 최첨단 강화 학습 기반 모델에 대한 소규모 실험 결과를 제시합니다. 실험 결과, 양 모델 모두에서 의견에 동조하는 경향과 논쟁적인 가치에 대한 낮은 수정 품질이 동시에 나타나는 것을 확인했습니다. PRS는 다원주의를 위한 상호작용적 전제 조건(눈에 띄는 의견 불일치, 원칙에 기반한 수정)을 측정하는 것이며, 완전한 다원주의 자체를 측정하는 것은 아닙니다. 우리는 이러한 차이점을 논의하고, '원칙'의 기준이 누구의 기준인지에 대한 성찰적인 질문을 진지하게 다루며, 다원주의가 가장 결정적으로 형성되거나 파괴되는 것은 배포 및 거버넌스 단계, 즉 인터페이스, 선호도 데이터 파이프라인, 그리고 감사 인프라에서 이루어진다는 주장을 펼칩니다.
Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or proportionally represent (Distributional) diverse human values. We argue that aggregation alone is an incomplete primitive for deployed pluralistic alignment. Under genuine value pluralism, the failure mode of contemporary RLHF-trained assistants is not insufficient coverage but sycophantic consensus: a learned tendency to agree with, validate, and minimise friction with the immediate interlocutor. Because deployed AI systems now mediate consequential deliberation across health, civic life, labour, and governance, the collapse of disagreement at the interaction layer is not a narrow technical concern but a structural failure with distributive consequences. We reframe pluralistic alignment around three conversational mechanisms drawn from Grice's maxims: scoping (acknowledging the limits of one's perspective), signalling (surfacing value-conflict rather than smoothing it over), and repair (revising one's position on principled grounds, not on user pressure). We formalise a metric, the Pluralistic Repair Score (PRS), distinguishing principled revision from capitulation, and present a small-scale empirical illustration on two frontier RLHF-trained models (Claude Sonnet 4.5, N=198; GPT-4o, N=100) showing that, for both, agreement-following coexists with low repair-quality on contested-value prompts. PRS measures an interactional precondition for pluralism (visible disagreement; principled revision) rather than pluralism in full; we discuss the difference, take seriously the reflexive question of whose "principled" counts, and argue that pluralism is most decisively made or unmade at the deployment-governance layer: interfaces, preference-data pipelines, and audit infrastructure.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.