DelusionEval: AI 챗봇에서 망상 관련 행동 측정
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
정신 건강 전문가들은 대규모 언어 모델(LLM)과의 상호 작용으로 인해 발생할 수 있는 심리적 위험, 특히 인간과 LLM의 문제가 되는 행동이 시간이 지남에 따라 서로 강화되는 '망상 고리' 현상에 대해 우려를 표명해 왔습니다. LLM 기반 챗봇의 사용이 증가함에 따라, 실제 사용자에게서 경험되는 심리적 피해 사례를 기반으로 한 평가 시스템 구축이 시급합니다. 본 연구에서는 사용자의 망상을 조장하는 행동을 나타내는 경향성을 검증하기 위한 평가 프로토콜인 DelusionEval을 개발했습니다. 우리는 18명의 참가자로부터 수집된 589개의 고유한 대화 기록(총 12,591건의 메시지)을 사용하여 각 모델을 평가했습니다. 이 데이터는 망상 및 심리적 피해를 경험한 사용자들의 메시지로 구성되었습니다. 분석 결과, 평가된 LLM이 망상 관련 행동을 나타내는 경향성은 모델 크기, 출시일 또는 테스트 시 추론 기능 유무와 일관되게 연관되지 않았습니다. 그러나 이전 메시지의 컨텍스트를 확장하면 망상 관련 행동의 발생률이 크게 증가하며, 이는 LLM 안전성 평가에서 컨텍스트의 중요성을 보여주는 증거입니다. 예를 들어, 사용자가 자살 충동을 표현할 때, 자살 시도를 억제하지 못하는 비율은 대화 기록 앞에 추가되는 메시지가 350건일 때 30.0%에서 41.1%로 증가했습니다. GPT, Claude 등 모든 모델 계열에서 상당한 수준의 망상 관련 행동이 나타났습니다. 계열 내에서도 후기 모델, 더 큰 모델 또는 더 높은 추론 능력을 가진 모델이 모든 행동 범주에서 일관되게 더 나은 성능을 보이는 것은 아닙니다. 이러한 결과는 LLM이 미칠 수 있는 잠재적인 심리적 영향에 대한 우려를 제기하며, 실제 인간-AI 상호 작용에 대한 더욱 엄격한 연구의 필요성을 강조합니다.
Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which concerning human and LLM behaviors reinforce each other over time. With growing public use of LLM-powered chatbots, there is an urgent need to build evaluations grounded in real-world episodes of psychological harm experienced by users. We developed DelusionEval, an evaluation protocol that tests a model's tendencies to exhibit behaviors linked to promoting user delusions. We prompt each model with 589 unique conversation histories from 18 participants, comprising 12,591 messages from users who experienced delusions and psychological harm. We find that the tendency of an evaluated LLM to exhibit delusion-linked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. However, extending the context of prior messages substantially increases rates of delusion-linked behaviors, providing evidence for the importance of context in LLM safety evaluation. For example, the rate of failing to discourage self-harm when the user expresses suicidal ideation increases from 30.0% to 41.1% when an additional 350 messages are prepended to the conversation history. All model families (e.g., GPT, Claude) exhibit substantial rates of delusion-linked behaviors. Within families, later, larger, or higher-reasoning models are not uniformly better across all behavior categories. Our results raise concerns regarding the potential psychological impact of LLMs and the need for more rigorous studies of real-world human-AI interaction.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.