AI 콘텐츠 검열이 치료 상담에 미치는 영향
AI Content Moderation in Therapy Conversations
최근 대규모 언어 모델(LLM)은 감정적 지원을 제공하는 데 점점 더 많이 활용되고 있으며, 공식적인 치료 목적으로도 개발되고 있습니다. 그러나 ChatGPT나 Llama와 같은 LLM은 책임 및 안전상의 이유로 콘텐츠 검열 장치를 내장하고 있어 사용자와의 민감한 주제에 대한 논의를 제한하는 경우가 많습니다. 이러한 제약은 LLM이 치료사로서의 역할을 수행하는 능력에 영향을 미칠 수 있습니다. 본 연구에서는 OpenAI의 moderation endpoint, Meta의 Llama Guard, Google의 Shield Gemma 등 최첨단 콘텐츠 검열 시스템 3가지에 대해 알고리즘 감사를 실시하여 실제 치료 상담 내용이 얼마나 자주 부적절한 콘텐츠로 분류되는지 조사했습니다. 결과는 사용자와 조직이 LLM을 치료사 역할을 수행하도록 설계할 때 발생할 수 있는 제한 사항에 대한 시사점을 제공합니다.
Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LLMs like ChaptGPT or Llama are often developed with content moderation guardrails that prevent them from discussing sensitive subjects with users for both liability and safety purposes, and this inability to broach these subjects may affect their capacity as therapists. In this study, we perform an algorithm audit on three state-of-the-art moderation systems (OpenAI's moderation endpoint, Meta's Llama Guard, and Google's Shield Gemma) to investigate the extent to which these systems flag the content of real-life therapy sessions as undesirable. Our results raise implications for the limitations that users and organizations may encounter when designing LLMs to play the part of a therapist.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.