FanarGuard: 아랍어 언어 모델을 위한 문화적 맥락 인지 콘텐츠 필터
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
콘텐츠 필터는 언어 모델의 안전성 확보에 필수적인 요소입니다. 그러나 현재 대부분의 필터는 일반적인 안전성에만 초점을 맞추고 있으며, 문화적 맥락을 고려하지 못하는 경우가 많습니다. 본 연구에서는 아랍어와 영어를 모두 지원하며, 안전성과 문화적 적합성을 동시에 평가하는 이중 언어 콘텐츠 필터인 FanarGuard를 소개합니다. 인공 데이터 및 공개 데이터를 활용하여 46만 개 이상의 프롬프트-응답 쌍으로 구성된 데이터셋을 구축하고, LLM 심판단을 통해 무해성 및 문화적 인식 측면에서 점수를 매겼습니다. 이 데이터셋을 사용하여 두 가지 필터 변형을 학습시켰습니다. 또한, 아랍어 문화적 맥락에 대한 첫 번째 벤치마크를 개발하여, 1천 개 이상의 민감한 프롬프트-응답 쌍으로 구성되었으며, LLM이 생성한 응답은 인간 평가자가 주관적으로 평가했습니다. 실험 결과, FanarGuard는 안전성 벤치마크에서 최첨단 필터와 동등한 성능을 보이면서도, 인간 평가 기준과의 일치도가 평가자 간 신뢰도보다 높았습니다. 이러한 결과는 콘텐츠 필터에 문화적 인식 요소를 통합하는 것의 중요성을 강조하며, FanarGuard를 보다 맥락에 민감한 안전 장치를 구축하기 위한 실질적인 단계로 제시합니다.
Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on general safety and overlook cultural context. In this work, we introduce FanarGuard, a bilingual moderation filter that evaluates both safety and cultural alignment in Arabic and English. We construct a dataset of over 468K prompt and response pairs, drawn from synthetic and public datasets, scored by a panel of LLM judges on harmlessness and cultural awareness, and use it to train two filter variants. To rigorously evaluate cultural alignment, we further develop the first benchmark targeting Arabic cultural contexts, comprising over 1k norm-sensitive prompts with LLM-generated responses annotated by human raters. Results show that FanarGuard achieves stronger agreement with human annotations than inter-annotator reliability, while matching the performance of state-of-the-art filters on safety benchmarks. These findings highlight the importance of integrating cultural awareness into moderation and establish FanarGuard as a practical step toward more context-sensitive safeguards.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.