2605.27288v1 May 26, 2026 cs.CL

언제나 아첨은 아니다: 인식적 불확실성을 함수로 하는 LLM의 순응도 측정

It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

Juming Xiong
Juming Xiong
Citations: 128
h-index: 6
Kevin H. Guo
Kevin H. Guo
Citations: 3
h-index: 1
Avinash Baidya
Avinash Baidya
Citations: 77
h-index: 5
Katherine E. Brown
Katherine E. Brown
Citations: 54
h-index: 3
Zhijun Yin
Zhijun Yin
Citations: 240
h-index: 8
Bradley Malin
Bradley Malin
Citations: 41
h-index: 3
Xiang Gao
Xiang Gao
Citations: 89
h-index: 5
Chao Yan
Chao Yan
Citations: 75
h-index: 4

대규모 언어 모델(LLM)은 사용자의 반박에 대응하여 초기 입장을 포기하는 경향이 있습니다. 기존 연구에서는 이러한 현상을 인간 피드백을 통한 강화 학습 과정에서 학습된 아첨으로 간주했지만, 본 연구에서는 순응성이 모델의 추론 시점에서의 인식적 불확실성에 의해 더욱 크게 영향을 받는다고 가정합니다. 본 논문에서는 LLM의 순응 행동을 유발하는 메커니즘을 분리하기 위한 두 단계 평가 프레임워크인 MUSE를 소개합니다. 구체적으로, MUSE는 모델이 특정 질문에 대해 보이는 인식적 불확실성을, 후속 턴에서 사용자의 반박에 얼마나 쉽게 동조하는지 여부와 연결하여 분석합니다. 우리는 순응 행동을 유발하는 메커니즘이 아첨뿐만 아니라 다른 요인도 포함한다는 것을 보여줍니다. 특히, 우리는 순응 행동을 공동으로 유발하는 두 가지 상이한 요인을 규명했습니다. 첫째는 모델이 초기 응답에 대해 절대적인 확신을 가지고 있음에도 사용자의 반박에 동조하는 '아첨적 순응'이며, 둘째는 모델의 불확실성이 증가함에 따라 순응 가능성도 함께 증가하는 '불확실성 기반 순응'입니다. 또한, ablation study를 통해 아첨적 순응과 불확실성 기반 순응 모두가 1) 사용자가 LLM을 인지하는 전문성과 2) 사용자 제안의 타당성이 높아질수록 증가한다는 것을 입증했습니다. 더 나아가, MUSE는 alignment-induced 아첨(alignment에 의해 유발된 아첨)과 training-corpora-driven 불확실성(학습 데이터에 의한 불확실성)을 구별함으로써 보다 효과적인 개입 전략 개발에 기여합니다.

Original Abstract

Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behavior to sycophancy learned during reinforcement learning from human feedback, we hypothesize that conformity is also driven by a model's epistemic uncertainty at inference time. In this paper, we introduce MUSE, a two-stage evaluation framework to disentangle the mechanisms driving LLM conformity. Specifically, MUSE maps a model's epistemic uncertainty in responding to a query against its likelihood to yield to user pushback in a subsequent turn. We demonstrate that the mechanisms driving conformity extend beyond sycophancy alone. Specifically, we characterize two distinct factors that jointly drive conformity: sycophantic conformity, where a model aligns with user pushback even with absolute certainty in its initial response, and uncertainty-driven conformity, where a model's likelihood for conformity increases alongside its uncertainty. Furthermore, we conduct ablation studies to demonstrate that both sycophantic conformity and uncertainty-driven conformity grow with 1) the LLM's perceived expertise of the user and 2) the plausibility of the user's suggestions. More broadly, MUSE informs more targeted intervention strategies by distinguishing alignment-induced sycophancy and training-corpora-driven uncertainty.

1 Citations
0 Influential
4 Altmetric
21.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!