유해한 AI 아첨 현상의 측정 및 탐지
Measuring and Detecting Harmful AI Sycophancy
아첨적인 답변이 대규모 언어 모델(LLM)에서 점점 더 보편화되고 있으며, 기존 연구에서는 이러한 답변 중 일부가 유해할 수 있다는 점을 지적했습니다. 본 논문은 '선호도 기반 태도 전환 아첨(Preference-Induced Stance Reversal Sycophancy, PSRS)'이라는 특정 유형의 유해한 아첨 현상에 초점을 맞춥니다. PSRS는 모델이 사용자의 선호도에 부합하기 위해 초기 입장을 단순히 뒤바꾸는 경우를 의미합니다. 기존 연구가 주로 모델의 아첨 정도를 측정하는 데 집중했다면, 본 연구는 PSRS를 단일 답변에서 자동으로 탐지할 수 있는지 여부를 더 나아가 질문합니다. 이를 대규모로 조사하기 위해, 우리는 레이블이 지정된 PSRS 데이터를 수집하는 프레임워크인 CAP(Contrastive Anchor Probing)을 소개합니다. CAP를 17개의 공개 및 비공개 LLM에 적용하여, 12가지 일상적인 조언 분야에서 총 290,460개의 레이블이 지정된 답변을 수집했습니다. 본 연구는 세 가지 주요 질문을 중심으로 구성됩니다. (1) PSRS가 얼마나 자주 발생하는가? (2) PSRS를 얼마나 잘 탐지할 수 있는가? (3) 탐지가 새로운 모델로 얼마나 잘 일반화되는가? 먼저, PSRS 발생률이 LLM에 따라 5%에서 56%까지 다양하며, 성능이 더 뛰어난 모델일수록 아첨 현상이 덜 나타나는 것으로 확인되었습니다. 또한, PSRS를 답변 텍스트만으로 탐지하는 것이 가능하며, 탐지기는 훈련 데이터를 통해 미묘한 PSRS 패턴을 학습해야 함을 보여주었습니다. 새로운 LLM이 빠르게 등장함에 따라, 탐지기는 필연적으로 이전에 보지 못한 모델과 마주하게 되므로, 모델 간 일반화는 중요한 연구 목표입니다. 우리는 탐지 성능이 이전에 보지 못한 모델에서 저하되는 것을 확인하고, 이러한 문제를 해결하기 위한 초기 접근 방식을 제안합니다. 본 연구에서 사용한 데이터셋 및 코드를 공개하여 향후 연구를 지원할 예정입니다.
Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user's stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also be detected automatically from a single response. To investigate this at scale, we introduce CAP (Contrastive Anchor Probing), a framework for collecting labeled PSRS data. Applying CAP to 17 open- and closed-source LLMs, we collect 290,460 labeled responses across 12 everyday-advice domains. We organize our study around three research questions. (1) How often does PSRS occur? (2) How well can it be detected? (3) How does detection generalize to unseen models? We first reveal that PSRS rates range from 5% to 56% across LLMs, with more capable models being less sycophantic. Next, we show that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data. Because new LLMs appear rapidly, detectors inevitably encounter unseen models, making cross-model generalization an important framework goal. We demonstrate that detection performance drops on unseen models and propose an initial approach to address this challenge. We will release our dataset and code to support future research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.