2601.06700v1 Jan 10, 2026 cs.CL

생성형 거대 언어 모델의 유해성 분석

Characterising Toxicity in Generative Large Language Models

Yuhang Wu
Yuhang Wu
Citations: 250
h-index: 4
Zhiyao Zhang
Zhiyao Zhang
Citations: 13
h-index: 2
Yazan Mash'al
Yazan Mash'al
Citations: 0
h-index: 0

최근 몇 년 동안 어텐션 메커니즘의 발전은 자연어 처리(NLP) 분야를 크게 발전시켰으며, 텍스트 처리 및 텍스트 생성 방식을 혁신했습니다. 이는 트랜스포머 기반 디코더 전용 아키텍처를 통해 이루어졌으며, 뛰어난 텍스트 처리 및 생성 능력을 바탕으로 NLP 분야에서 널리 사용되고 있습니다. 이러한 획기적인 발전에도 불구하고, 언어 모델(LM)은 여전히 부적절하거나 공격적이며 유해한 응답과 같은 원치 않는 결과를 생성할 수 있습니다. 이러한 결과를 통칭하여 ``유해한`` 결과라고 부릅니다. 강화 학습 기반 인간 피드백(RLHF)과 같은 방법이 모델의 출력을 인간의 가치에 맞추기 위해 개발되었지만, 이러한 안전장치는 신중하게 작성된 프롬프트를 통해 종종 우회될 수 있습니다. 따라서 본 논문에서는 LLM이 프롬프트를 받았을 때 얼마나 자주 유해한 콘텐츠를 생성하는지, 그리고 생성 모델에서 그러한 결과의 생성에 영향을 미치는 어휘적 및 구문적 요인을 분석합니다.

Original Abstract

In recent years, the advent of the attention mechanism has significantly advanced the field of natural language processing (NLP), revolutionizing text processing and text generation. This has come about through transformer-based decoder-only architectures, which have become ubiquitous in NLP due to their impressive text processing and generation capabilities. Despite these breakthroughs, language models (LMs) remain susceptible to generating undesired outputs: inappropriate, offensive, or otherwise harmful responses. We will collectively refer to these as ``toxic'' outputs. Although methods like reinforcement learning from human feedback (RLHF) have been developed to align model outputs with human values, these safeguards can often be circumvented through carefully crafted prompts. Therefore, this paper examines the extent to which LLMs generate toxic content when prompted, as well as the linguistic factors -- both lexical and syntactic -- that influence the production of such outputs in generative models.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!