2601.03121v1 Jan 06, 2026 cs.CL

ToxiGAN: LLM 기반 방향성 적대적 생성을 통한 유해 데이터 증강

ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation

Peiran Li
Peiran Li
Citations: 3,461
h-index: 4
Jan Fillies
Jan Fillies
Citations: 56
h-index: 5
Adrian Paschke
Adrian Paschke
Citations: 78
h-index: 4

유해 언어 데이터의 제어 가능하고 클래스별 증강은 독성 분류의 견고성을 향상시키는 데 매우 중요하지만, 제한적인 감독과 데이터 분포의 편향으로 인해 어려운 과제입니다. 본 논문에서는 적대적 생성을 의미론적 지침과 결합하여 클래스 정보를 고려한 텍스트 증강 프레임워크인 ToxiGAN을 제안합니다. ToxiGAN은 GAN 기반 증강에서 흔히 발생하는 모드 붕괴 및 의미론적 편향 문제를 해결하기 위해, 두 단계의 방향성 학습 전략을 도입하고, 대규모 언어 모델(LLM)에서 생성된 중립적인 텍스트를 의미론적 균형추로 활용합니다. 기존 연구에서 LLM을 정적인 생성기로 취급하는 것과는 달리, 저희의 접근 방식은 균형 잡힌 지침을 제공하기 위해 중립적인 예시를 동적으로 선택합니다. 유해 샘플은 이러한 예시와 명시적으로 분리되도록 최적화되어, 클래스별 대비 신호를 강화합니다. 네 가지 혐오 발언 벤치마크에서의 실험 결과, ToxiGAN은 매크로-F1 점수와 혐오-F1 점수 모두에서 가장 뛰어난 평균 성능을 보였으며, 기존의 전통적인 증강 방법 및 LLM 기반 증강 방법보다 일관되게 우수한 성능을 나타냈습니다. 추가 분석 결과, 의미론적 균형추와 방향성 학습이 분류기의 견고성을 향상시키는 데 기여한다는 것을 확인했습니다.

Original Abstract

Augmenting toxic language data in a controllable and class-specific manner is crucial for improving robustness in toxicity classification, yet remains challenging due to limited supervision and distributional skew. We propose ToxiGAN, a class-aware text augmentation framework that combines adversarial generation with semantic guidance from large language models (LLMs). To address common issues in GAN-based augmentation such as mode collapse and semantic drift, ToxiGAN introduces a two-step directional training strategy and leverages LLM-generated neutral texts as semantic ballast. Unlike prior work that treats LLMs as static generators, our approach dynamically selects neutral exemplars to provide balanced guidance. Toxic samples are explicitly optimized to diverge from these exemplars, reinforcing class-specific contrastive signals. Experiments on four hate speech benchmarks show that ToxiGAN achieves the strongest average performance in both macro-F1 and hate-F1, consistently outperforming traditional and LLM-based augmentation methods. Ablation and sensitivity analyses further confirm the benefits of semantic ballast and directional training in enhancing classifier robustness.

1 Citations
1 Influential
2.5 Altmetric
15.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!