속임수를 간파하다: 순서 기반 추론 체인 정규화를 통한 끊임없이 변화하는 유해 채팅 대화 탐지
Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization
유해한 채팅 대화는 유형 변화 및 어휘 회피를 통해 끊임없이 변화하지만, 우리는 이러한 대화가 반복적인 주제, 유해 언어 지표, 심각도 계층 구조, 유형 특징 등 불변하는 원칙, 즉 '순서 기반 추론 체인(ORC)'을 공유한다는 것을 발견했습니다. 이는 빈번하게 변하는 어휘 표현에서 핵심 정보를 포착하는 데 도움이 될 수 있습니다. 우리는 ORC를 4개의 미분 가능한 단계(주제 -> 지표 -> 심각도 -> 유형)로 인코딩하고, 중간 수준의 감독 학습을 활용하여 구조적 정규화 역할을 하는 BRACE 모델을 제안합니다. 또한 프로토타입 기반 특징 증강 및 특징 경로 분리 기법을 통해 성능을 향상시켰습니다. 평가 결과, 4개 도메인 및 5가지 유해 범주에 대해 BRACE는 0.934의 유해 유형 매크로 F1 점수(RoBERTa-wwm-ext, 3회 평균)를 달성했으며, 디코더 백본(Qwen3-1.7B LoRA)을 사용할 경우 0.949까지 성능 향상을 보였습니다. 분석 결과, 모든 구성 요소가 BRACE의 성능에 기여하며, ORC의 구조적 분해는 의미적으로 모호한 유해 유형을 구별하는 데 도움이 됩니다. 면책 조항: 본 논문에는 일부 독자에게 불쾌감을 줄 수 있는 내용이 포함되어 있을 수 있습니다.
Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring topics, harm language indicators, severity hierarchies, and type characteristics, which can help us capture the key information in the frequently changing lexical expressions. We propose BRACE, which encodes the ORC as four differentiable stages (Topic -> Indicator -> Severity -> Type) with intermediate supervision, serving as a structured regularizer blended with direct heads, and supported by prototype-based feature augmentation and feature path disentanglement. The evaluation results show that, across 4 domains and 5 harm categories, BRACE achieves harm-type macro F1 of 0.934 (RoBERTa-wwm-ext, 3-seed mean), with decoder backbones (Qwen3-1.7B LoRA) reaching 0.949. Ablation studies show that all components contribute to BRACE, and the structural decomposition of ORC enables BRACE to distinguish harmful types with semantic ambiguity. Disclaimer: This paper may contain content that is disturbing to some readers.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.