SEAHateCheck: 동남아시아 저자원 언어에서 혐오 표현 탐지를 위한 기능 테스트
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia
혐오 표현 탐지는 주로 영어 및 중국어와 같은 고자원 언어에 의존하는 언어 자원에 크게 의존하며, 이는 동남아시아의 다양한 사회-언어적 맥락 속에서 온라인 혐오 표현 관리를 위한 도구를 개발하는 연구자와 플랫폼에 장벽을 만듭니다. 이를 해결하기 위해, 우리는 인도네시아, 태국, 필리핀, 베트남을 대상으로 인도네시아어, 타갈로그어, 태국어, 베트남어를 포함하는 데이터셋인 SEAHateCheck를 소개합니다. SEAHateCheck는 HateCheck의 기능 테스트 프레임워크를 기반으로 하며, SGHateCheck의 방법을 개선하여 문화적으로 관련성이 높은 테스트 케이스를 제공하며, 대규모 언어 모델을 활용하고 현지 전문가의 검증을 통해 정확성을 높였습니다. 최첨단 및 다국어 모델을 사용한 실험 결과, 특정 저자원 언어에서 혐오 표현 탐지에 한계가 있음을 확인했습니다. 특히, 타갈로그어 테스트 케이스의 모델 정확도가 가장 낮았는데, 이는 언어적 복잡성과 제한된 학습 데이터 때문일 가능성이 높습니다. 반면, 은어 기반의 기능 테스트는 모델이 문화적으로 미묘한 표현을 이해하는 데 어려움을 겪으면서 가장 어려운 것으로 나타났습니다. SEAHateCheck의 분석 결과는 모델의 암묵적인 혐오 표현 탐지 능력 부족과 반박 표현에 대한 모델의 어려움을 더욱 명확하게 보여줍니다. 본 연구는 이러한 동남아시아 언어들을 위한 최초의 기능 테스트 모음이며, 연구자들에게 견고한 벤치마크를 제공하여 포괄적인 온라인 콘텐츠 관리를 위한 실용적이고 문화적으로 적합한 혐오 표현 탐지 도구 개발을 촉진합니다.
Hate speech detection relies heavily on linguistic resources, which are primarily available in high-resource languages such as English and Chinese, creating barriers for researchers and platforms developing tools for low-resource languages in Southeast Asia, where diverse socio-linguistic contexts complicate online hate moderation. To address this, we introduce SEAHateCheck, a pioneering dataset tailored to Indonesia, Thailand, the Philippines, and Vietnam, covering Indonesian, Tagalog, Thai, and Vietnamese. Building on HateCheck's functional testing framework and refining SGHateCheck's methods, SEAHateCheck provides culturally relevant test cases, augmented by large language models and validated by local experts for accuracy. Experiments with state-of-the-art and multilingual models revealed limitations in detecting hate speech in specific low-resource languages. In particular, Tagalog test cases showed the lowest model accuracy, likely due to linguistic complexity and limited training data. In contrast, slang-based functional tests proved the hardest, as models struggled with culturally nuanced expressions. The diagnostic insights of SEAHateCheck further exposed model weaknesses in implicit hate detection and models' struggles with counter-speech expression. As the first functional test suite for these Southeast Asian languages, this work equips researchers with a robust benchmark, advancing the development of practical, culturally attuned hate speech detection tools for inclusive online content moderation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.