본 논문에서는 Shieldstral을 소개합니다. Shieldstral은 30억 개의 파라미터를 가진 정책 적응형 다중 모드 안전 분류기로, 텍스트 안전성 평가에서 거의 7배 더 큰 모델과 동등하거나 뛰어난 성능을 보이며, 다중 모드 안전 분류 분야에서 새로운 최고 수준의 성능을 달성합니다. Shieldstral은 콘텐츠 검열을 이진 질문-답변 문제로 정의합니다. 이러한 간단한 접근 방식을 통해 다양한 검열 작업을 하나의 '예/아니오' 문제로 통합하여, 서로 다른 분류 체계를 가진 다양한 안전 데이터 세트를 단일 학습 프레임워크 내에서 통합할 수 있습니다. 본 논문에서는 약 5억 4천만 개의 샘플을 포함하는 데이터 구축 방법을 제시하며, 정책 적응성을 평가하기 위한 상세한 평가 세트도 제공합니다. 이러한 요소들이 결합되어, 작은 크기의 적응형 모델이 훨씬 더 큰 모델과 동등하거나 뛰어난 성능을 발휘할 수 있도록 합니다.
Original
Abstract
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.