2607.29585v1 Jul 31, 2026 cs.CL

아첨(Sycophancy)이 협력적 시각-언어 작업에서 인지적 경계심을 저해한다

Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks

C. Bonial
C. Bonial
Citations: 2,813
h-index: 20
Rachel Rudinger
Rachel Rudinger
Citations: 2
h-index: 1
Rupak Sarkar
Rupak Sarkar
Citations: 235
h-index: 10
Neha Srikanth
Neha Srikanth
Citations: 140
h-index: 6
Saloni Gupta
Saloni Gupta
Citations: 467
h-index: 2
Philip Resnik
Philip Resnik
Citations: 14
h-index: 2

협력적인 대화에서, 사람들은 상대방이 새로운 정보를 공유함에 따라 자신의 믿음을 반복적으로 업데이트합니다. 인지적 경계심이 높은 참여자는 새로운 정보가 기존의 신념과 충돌할 때 이를 감지하고 이러한 갈등을 해결하기 위한 조치를 취합니다. AI 시스템이 복잡한 협력 작업에서 신뢰할 수 있는 파트너 역할을 수행하려면, 마찬가지로 들어오는 정보를 자체적인 증거와 공유된 맥락에 비추어 평가해야 하며, 불일치가 발생하면 이를 적절하게 반영해야 합니다. 본 연구에서는 협력적 환경에서 시각-언어 모델의 인지적 경계심을 측정하기 위해, 정보 비대칭적인 대화 기반 '차이점 찾기' 과제를 제시합니다. 두 개의 모델은 각각 다른 이미지를 개인적으로 보여주고, 대화를 통해 이미지들이 동일한지 여부를 판단하거나, 동일하지 않은 경우 차이점을 식별해야 합니다. 그러나 모델들은 이 과제에서 일관되게 실패합니다. 그들은 종종 자신의 이미지에 포함된 중요한 증거를 간과하고, 정당하지 않더라도 대화 상대방의 의견에 동조하는 경향을 보입니다. 이러한 인지적 경계심 위반 현상을 아첨(sycophancy)이라는 더 광범위한 행동 패턴과 연결합니다. 아첨은 협력적인 목표 지향적 대화에서 과도한 양보와 취약한 증거 기반으로 나타납니다. 본 연구의 결과는, 작업에 독립적인 아첨 예제로부터 학습된 벡터를 사용하여 모델의 아첨 경향을 줄이면 인지적 경계심과 관련된 오류를 감소시킬 수 있으며, 이는 모델이 자신의 증거를 더욱 정확하게 보고하고, 궁극적으로 정보 비대칭적인 협력 작업에서 더 신뢰할 수 있는 파트너가 되도록 하는 데 기여한다는 것을 보여줍니다.

Original Abstract

To maintain common ground in cooperative conversation, humans iteratively update their beliefs as conversation participants share new information; participants who are epistemically vigilant detect when new information conflicts with prior beliefs and take steps to repair these conflicts. In order for AI systems to serve as reliable partners in complex cooperative tasks, they must similarly weigh incoming information against their own private evidence and shared context and appropriately surface inconsistencies when they arise. To measure the epistemic vigilance of vision-language models in cooperative settings, we present an information-asymmetric, dialog-based "spot-the-difference" task. Two models are privately shown one image each, and must determine through conversation whether the images are identical or, if not, identify the difference. Models routinely fail at this: they frequently overlook key evidence in their private image in favor of agreeing with their conversational partner, even when their agreement is unwarranted. We relate these violations of epistemic vigilance to the broader behavior of sycophancy, which manifests itself in cooperative goal-oriented dialog as over-accommodation and weak evidential grounding. Our results show that model steering to reduce sycophancy with a vector learned from task-agnostic sycophancy examples can reduce epistemic vigilance-related errors, making models more faithful reporters of their evidence, and in turn, more reliable partners in information-asymmetric cooperative tasks.

1 Citations
0 Influential
10 Altmetric
51.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!