2607.18114v1 Jul 20, 2026 cs.CL

정렬 미세 조정이 LLM에서 아첨 및 관련 신호 유발 편향의 표현에 어떤 영향을 미치는가?

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

Zhijing Jin
Zhijing Jin
Citations: 2
h-index: 1
Florent Draye
Florent Draye
Citations: 16
h-index: 3
Terry Jingchen Zhang
Terry Jingchen Zhang
Citations: 60
h-index: 4
Bernhard Schölkopf
Bernhard Schölkopf
Citations: 999
h-index: 18
Prakhar Gupta
Prakhar Gupta
Carnegie Mellon University
Citations: 7,503
h-index: 14

최신 LLM은 놀라울 정도로 단순한 입력 프롬프트의 변화에 매우 취약합니다. 사소한 힌트, 잘못된 레이블이 붙은 소량 예제 또는 가짜 이전 어시스턴트 대화는 원래 올바른 답변을 뒤집을 수 있습니다. 본 연구에서는 이러한 취약성이, 즉 아첨 및 관련 신호 유발 편향이 모델 내에서 어디에 존재하는지 분석합니다. 5개의 모델 패밀리와 7가지 BCT 편향 유형을 대상으로, 은닉 상태로부터 각 편향 방향을 추출하고 세 가지 방법(프롬프트 기반 분석, 데이터셋 하나를 제외한 전이 학습, 인과적 개입)을 통해 이를 정량화했습니다. 이러한 취약성은 주로 사전 훈련보다는 정렬 미세 조정에 의해 발생합니다. 사전 훈련된 기본 모델은 이러한 편향에 거의 영향을 받지 않으며, 이들의 활성화 값은 질문 내용 외에는 특정 신호와 관련된 정보를 담고 있지 않습니다. 정렬된 모델 내에서 각 편향은 일관된 방향을 가지며, 우리는 이를 해독하고 제어하여 테스트한 모든 패밀리에서 편향되지 않은 답변을 얻을 수 있습니다. 그러나 이러한 편향들은 표현적으로 구별됩니다. 교차 편향 간의 상호 작용은 편향 범주 자체의 속성이라기보다는 모델에 특이하며, 심지어 행동적으로 유사한 편향조차도 서로 다른 방향을 가집니다. 동일한 개입 방법은 또한 온건한 비편향화 도구로 사용될 수 있으며, 모든 지시형 모델에서 대부분의 올바른 답변을 유지하면서 편향으로 인한 오류를 상당 부분 수정할 수 있습니다. 따라서 신호 유발 편향은 LLM의 단일 결함이 아니라, 정렬 미세 조정에 의해 설치되는 다양한 종류의 상호 연관된 방향들의 집합으로 이해하는 것이 가장 적절합니다.

Original Abstract

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven BCT bias types, we extract a per-bias direction from hidden states and triangulate it through three measures: probing, leave-one-dataset-out transfer, and causal intervention. The susceptibility is largely installed by alignment tuning rather than pretraining: pretrained base models barely cave to these biases, and their activations carry no cue-specific signal beyond question content. Within aligned models, each bias becomes a single coherent direction that we can both decode and steer along, recovering the unbiased answer across every family we test. The biases stay representationally distinct, however: cross-bias entanglement is model-specific rather than a property of the bias category, and even behaviorally similar biases occupy different directions. The same intervention also serves as a modest debiasing tool, recovering a meaningful share of bias-induced errors while preserving most correct answers across all instruct families. Cue-induced bias is therefore best understood not as a single flaw in LLMs but as a family of distinct, causally active directions that alignment tuning installs.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!