ContiGuard: 진화하는 회피적 변형에 대한 지속적인 유해 콘텐츠 탐지 프레임워크
ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
유해 콘텐츠 탐지는 온라인 소셜 환경에서 유해한 콘텐츠(예: 혐오 댓글, 게시물, 메시지)의 확산을 줄여 건강한 온라인 환경을 유지하는 데 중요한 역할을 합니다. 그러나 악의적인 사용자는 지속적으로 유해 콘텐츠를 숨기고 탐지기를 회피하기 위한 회피적 변형을 개발합니다. 기존 탐지기 또는 방법은 시간이 지남에 따라 고정되어 있으며 이러한 진화하는 회피 기술에 효과적으로 대응하기 어렵습니다. 따라서 지속적인 학습은 진화하는 변형에 대한 탐지 능력을 동적으로 업데이트하는 데 적합한 접근 방식입니다. 그러나 다양한 변형으로 인해 탐지기가 변형된 텍스트에 대한 지속적인 학습이 어렵습니다. 더욱 중요한 것은 변형으로 인해 발생하는 노이즈가 의미를 왜곡하여 이해도를 저하시키고, 또한 중요한 특징 학습을 방해하여 탐지가 변형에 민감하게 반응하도록 만듭니다. 이러한 요인들은 진화하는 변형에 대한 지속적인 학습의 어려움을 더욱 가중시킵니다. 본 연구에서는 첫 번째로, 시간 경과에 따라 진화하는 변형된 텍스트(지속적인 유해 콘텐츠 탐지라고 함)에 대한 탐지기의 지속적인 학습을 위해 설계된 프레임워크인 ContiGuard를 제안합니다. ContiGuard는 탐지기가 지속적으로 기능을 업데이트하고 진화하는 변형에 대한 지속적인 견고성을 유지할 수 있도록 합니다. 구체적으로, 이해도를 높이기 위해 LLM(대규모 언어 모델) 기반의 의미 풍부화 전략을 제안합니다. 이 전략은 LLM이 추출한 가능한 의미와 유해성 관련 단서를 변형된 텍스트에 동적으로 통합하여 이해도를 향상시킵니다. 또한, 중요하지 않은 특징을 줄이고 중요한 특징을 강화하기 위해, 판별력 기반의 특징 학습 전략을 제안합니다. 이 전략은 판별력이 높은 특징을 강화하고 덜 판별적인 특징을 억제하여 탐지를 위한 견고한 분류 경계를 형성합니다.
Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social environment. However, malicious users persistently develop evasive perturbations to disguise toxic content and evade detectors. Traditional detectors or methods are static over time and are inadequate in addressing these evolving evasion tactics. Thus, continual learning emerges as a logical approach to dynamically update detection ability against evolving perturbations. Nevertheless, disparities across perturbations hinder the detector's continual learning on perturbed text. More importantly, perturbation-induced noises distort semantics to degrade comprehension and also impair critical feature learning to render detection sensitive to perturbations. These amplify the challenge of continual learning against evolving perturbations. In this work, we present ContiGuard, the first framework tailored for continual learning of the detector on time-evolving perturbed text (termed continual toxicity detection) to enable the detector to continually update capability and maintain sustained resilience against evolving perturbations. Specifically, to boost the comprehension, we present an LLM-powered semantic enriching strategy, where we dynamically incorporate possible meaning and toxicity-related clues excavated by LLM into the perturbed text to improve the comprehension. To mitigate non-critical features and amplify critical ones, we propose a discriminability-driven feature learning strategy, where we strengthen discriminative features while suppressing the less-discriminative ones to shape a robust classification boundary for detection...
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.