2505.19766v4 May 26, 2025 cs.CL

PAM: 정책에 부합하는 대규모 모더레이션 필터 학습

PAM: Training Policy-Aligned Moderation Filters at Scale

Enes Altinisik
Enes Altinisik
Citations: 265
h-index: 8
Masoomali Fatehkia
Masoomali Fatehkia
Citations: 514
h-index: 9
H. Sencar
H. Sencar
Citations: 4,637
h-index: 32
Mohamed Osman
Mohamed Osman
Citations: 31
h-index: 3

대규모 언어 모델(LLM)은 여전히 오정렬 및 우회 공격에 취약하며, 외부 안전장치인 모더레이션 필터가 필수적입니다. 그러나 기존 필터는 종종 안전에만 초점을 맞춰 실제 환경에서의 광범위한 정렬 요구사항을 충족하지 못합니다. 본 논문에서는 사용자 정의 정책 기반의 맞춤형 모더레이션 필터를 학습하는 유연한 프레임워크인 Policy Aligned Moderation (PAM)을 소개합니다. PAM은 인간이 작성한 예제에 의존하지 않고 자동으로 학습 데이터를 생성하여, 다양한 응용 분야별 정렬 목표 및 생성 정책에 대한 확장 가능한 지원을 가능하게 합니다. PAM으로 학습된 필터는 최첨단 안전 모더레이션 필터 및 정책 추론 모델과 동등한 성능을 보이며, 연령 제한, 식단 관련 요구 사항, 문화적 적합성, 의료 지침 제한 등을 목표로 하는 새로 도입된 사용자 주석 정책 준수 벤치마크인 PAMbench에서 더 뛰어난 성능을 나타냅니다. 이러한 성능 향상은 PAM 필터가 정책 기반 추론 모델보다 5~100배 빠른 속도로 작동하면서 달성됩니다.

Original Abstract

Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet existing filters often focus narrowly on safety, falling short of the broader alignment needs seen in real-world deployments. We introduce Policy Aligned Moderation (PAM), a flexible framework for training custom moderation filters grounded in user-defined policies that extend beyond conventional safety objectives. PAM automates training data generation without relying on human-written examples, enabling scalable support for diverse, application-specific alignment goals and generation policies. PAM-trained filters match the performance of state-of-the-art safety moderation filters and policy reasoning models, and outperform them on PAMbench, four newly introduced user-annotated policy enforcement benchmarks that target age restrictions, dietary accommodations, cultural alignment, and limitations in medical guidance. These performance gains are achieved while the PAM filter runs 5-100x faster at inference than policy-conditioned reasoning models.

1 Citations
0 Influential
16 Altmetric
81.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!