2604.05279v1 Apr 07, 2026 cs.AI

압력, 그게 뭔 소용 있나? 언어 모델에서의 아첨 현상 분리: 보상 분해를 통한 접근

Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition

Muhammad Ahmed Mohsin
Muhammad Ahmed Mohsin
Citations: 4
h-index: 1
Ahsan Bilal
Ahsan Bilal
Citations: 72
h-index: 6
Muhammad Umer
Muhammad Umer
Citations: 136
h-index: 8
Emily Fox
Emily Fox
Citations: 42
h-index: 2

대규모 언어 모델은 아첨이라는 경향을 보이며, 이는 증거와 관계없이 사용자의 선호나 권위 징후에 따라 모델이 자신의 입장을 바꾸는 현상입니다. 기존의 정렬 방법은 이 문제를 해결하는 데 실패하는데, 이는 스칼라 보상 모델이 두 가지 뚜렷한 실패 모드를 하나의 신호로 혼동하기 때문입니다. 첫 번째는 사회적 압력 하에서 올바른 답변을 수정하는 '압력 굴복', 두 번째는 제공된 맥락을 완전히 무시하는 '증거 무시'입니다. 본 연구에서는 압력 독립성과 증거 민감성을 형식적으로 정의하여 아첨 현상을 조작적으로 정의하고, 분리된 학습을 위한 작업 프레임워크를 제시합니다. 우리는 보상 분해를 통한 아첨 감소의 첫 번째 접근 방식을 제안하며, 5가지 구성 요소로 이루어진 그룹 상대 정책 최적화(GRPO) 보상을 도입합니다. 이 보상은 학습 신호를 압력 저항, 맥락 충실도, 입장 일관성, 동의 억제, 사실 정확성이라는 다섯 가지 항으로 분해합니다. 우리는 세 가지 권위 수준과 두 가지 대조적인 증거 맥락에서 압력이 없는 기준 데이터와 압력이 가해진 데이터를 쌍으로 묶어 학습을 진행합니다. 다섯 개의 기본 모델에 대해, 우리의 두 단계 파이프라인은 모든 측정 지표에서 일관되게 아첨 현상을 감소시키며, 각 보상 항이 독립적인 행동 차원을 제어한다는 것을 확인하는 분석을 수행했습니다. 학습된 압력 저항성은 우리의 학습 방법론 및 프롬프트 구조를 넘어 일반화되며, 학습 중에 그러한 압력 형태가 없었음에도 불구하고 SycophancyEval에서 최대 17포인트까지 답변 유도 아첨 현상을 감소시킵니다.

Original Abstract

Large language models exhibit sycophancy, the tendency to shift their stated positions toward perceived user preferences or authority cues regardless of evidence. Standard alignment methods fail to correct this because scalar reward models conflate two distinct failure modes into a single signal: pressure capitulation, where the model changes a correct answer under social pressure, and evidence blindness, where the model ignores the provided context entirely. We operationalise sycophancy through formal definitions of pressure independence and evidence responsiveness, serving as a working framework for disentangled training rather than a definitive characterisation of the phenomenon. We propose the first approach to sycophancy reduction via reward decomposition, introducing a multi-component Group Relative Policy Optimisation (GRPO) reward that decomposes the training signal into five terms: pressure resistance, context fidelity, position consistency, agreement suppression, and factual correctness. We train using a contrastive dataset pairing pressure-free baselines with pressured variants across three authority levels and two opposing evidence contexts. Across five base models, our two-phase pipeline consistently reduces sycophancy on all metric axes, with ablations confirming that each reward term governs an independent behavioural dimension. The learned resistance to pressure generalises beyond our training methodology and prompt structure, reducing answer-priming sycophancy by up to 17 points on SycophancyEval despite the absence of such pressure forms during training.

2 Citations
0 Influential
4 Altmetric
22.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!