2606.29887v1 Jun 29, 2026 cs.AI

SafePyramid: 문맥 내 정책 가이드레일링을 위한 계층적 벤치마크

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

Yuhao Sun
Yuhao Sun
Citations: 7
h-index: 2
Jiacheng Zhang
Jiacheng Zhang
Citations: 50
h-index: 2
Haoyu He
Haoyu He
Citations: 74
h-index: 5
Senjian Zhang
Senjian Zhang
Citations: 0
h-index: 0
Sheng-Ya Wang
Sheng-Ya Wang
Citations: 0
h-index: 0
Xiaolei Xu
Xiaolei Xu
Citations: 0
h-index: 0
Meng Shen
Meng Shen
Citations: 0
h-index: 0
Feng Liu
Feng Liu
Citations: 0
h-index: 0

실제 응용 프로그램에서 가이드레일은 미리 정의된 위험 분류 체계에 의존하는 것이 아니라, 애플리케이션별 안전 정책에 따라 사용자-모델 상호 작용의 잠재적인 위험을 식별하는 역할을 합니다. 본 연구에서는 문맥 내 정책 가이드레일링이라는 패러다임을 바탕으로 이러한 설정을 연구합니다. 여기서 가이드레일은 컨텍스트에 제공된 정책 사양을 기반으로 안전 위반 여부를 예측합니다. 이 기능을 체계적으로 평가하기 위해, 우리는 10개의 도메인에 걸쳐 총 1,000건의 다중 회화 데이터와 각 회화에 대응하는 3,000개의 애플리케이션별 정책으로 구성된 안전 벤치마크 'SafePyramid'를 소개합니다. SafePyramid는 개별 규칙 이해를 평가하는 L0, 규칙 간 의존성을 평가하는 L1, 그리고 컨텍스트 내에서 정의된 새로운 정책 프레임워크에 대한 적응력을 평가하는 L2의 세 가지 난이도 수준으로 평가를 구성합니다. 벤치마크의 품질을 보장하기 위해, 엄격한 다단계 파이프라인을 사용하여 벤치마크를 구축하고 검증했습니다. SafePyramid를 사용하여 10개의 최첨단 LLM(대규모 언어 모델)과 5개의 정책 구성 가능한 가이드레일을 평가한 결과, 문맥 내 정책 가이드레일링은 여전히 매우 어려운 과제라는 것을 확인했습니다. 가장 성능이 좋은 모델인 GPT-5.5조차도 L0, L1 및 L2 수준에서 각각 54.0%, 35.3% 및 12.9%의 경우에만 모든 위반 규칙을 정확하게 식별했습니다. 이러한 결과는 현재 가이드레일의 한계를 보여주며, 정책을 안정적으로 실행하고, 규칙 간 의존성을 해결하며, 새로운 정책 프레임워크에 적응할 수 있는 더욱 강력한 문맥 내 정책 가이드레일의 필요성을 강조합니다.

Original Abstract

In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies, rather than relying on predefined risk taxonomies. In this work, we study this setting under the paradigm of in-context policy guardrailing, where guardrails predict safety violations based on policy specifications provided in context. To systematically evaluate this capability, we introduce SafePyramid, a safety benchmark comprising 1,000 multi-turn conversations across 10 domains and 3,000 corresponding application-specific policies, which together contain 61,699 distinct natural-language rules. SafePyramid organizes the evaluation into three difficulty levels: L0 evaluates individual-rule understanding, L1 evaluates reasoning over rule dependencies, and L2 evaluates adaptation of full novel policy frameworks defined in context. To ensure benchmark quality, we employ a rigorous multi-stage pipeline to construct and validate the benchmark. Using SafePyramid, we evaluate 10 frontier LLMs and 5 policy-configurable guardrails and find that in-context policy guardrailing remains highly challenging: even the best-performing model, GPT-5.5, exactly identifies the full set of violated rules in only 54.0%, 35.3%, and 12.9% cases on L0, L1, and L2, respectively. These results highlight the limitations of current guardrails and call for stronger in-context policy guardrails that can reliably execute policies, resolve rule dependencies, and adapt to novel policy frameworks.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!