HateXScore: 혐오 발언 설명의 추론 품질을 평가하기 위한 지표 모음
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
혐오 발언 탐지는 콘텐츠 관리의 핵심 요소이지만, 현재의 평가 프레임워크는 텍스트가 왜 혐오적으로 간주되는지에 대한 이유를 거의 평가하지 않습니다. 본 논문에서는 모델 설명의 추론 품질을 평가하기 위해 설계된 네 가지 구성 요소로 이루어진 지표 모음인 HateXScore를 소개합니다. HateXScore는 (i) 결론의 명확성, (ii) 인용된 부분의 충실성과 인과적 근거, (iii) 보호 대상 집단 식별 (정책 구성 가능), (iv) 이러한 요소 간의 논리적 일관성을 평가합니다. HateXScore는 6개의 다양한 혐오 발언 데이터 세트를 대상으로 평가되었으며, 정확도 또는 F1과 같은 표준 지표로는 파악하기 어려운 해석 가능성 실패 및 주석 불일치를 진단하는 데 사용될 수 있습니다. 또한, 인간 평가 결과는 HateXScore와 높은 일관성을 보이며, 이는 HateXScore가 신뢰할 수 있고 투명한 관리를 위한 실용적인 도구임을 입증합니다. [경고: 본 논문에는 일부 독자에게 불쾌감을 줄 수 있는 민감한 내용이 포함되어 있습니다.]
Hateful speech detection is a key component of content moderation, yet current evaluation frameworks rarely assess why a text is deemed hateful. We introduce \textsf{HateXScore}, a four-component metric suite designed to evaluate the reasoning quality of model explanations. It assesses (i) conclusion explicitness, (ii) faithfulness and causal grounding of quoted spans, (iii) protected group identification (policy-configurable), and (iv) logical consistency among these elements. Evaluated on six diverse hate speech datasets, \textsf{HateXScore} is intended as a diagnostic complement to reveal interpretability failures and annotation inconsistencies that are invisible to standard metrics like Accuracy or F1. Moreover, human evaluation shows strong agreement with \textsf{HateXScore}, validating it as a practical tool for trustworthy and transparent moderation. \textcolor{red}{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.}
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.