2605.27851v1 May 27, 2026 cs.AI

맥락이 바뀌면 안전도 깨진다: 정렬된 언어 모델에서 취약한 안전성 진단

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

Dasol Choi
Dasol Choi
Citations: 5
h-index: 1
Alex Kwon
Alex Kwon
Citations: 0
h-index: 0

안전성 벤치마크 점수는 배포 준비 상태에 대한 불완전한 증거를 제공합니다. 정렬된 언어 모델은 종종 상황 변화로 인해 어떤 행동이 안전한지 바뀌더라도 엄격한 규칙을 따르는 경향이 있습니다. 우리는 이러한 현상을 '취약한 안전성'이라고 부릅니다. 이를 진단하기 위해, 우리는 컨텍스트-플립 평가 방법을 도입하여 12개의 모델을 안전성 벤치마크(PacifAIst) 및 두 가지 상식 제어 실험에서 테스트했습니다. 여기에는 기본적으로 안전한 행동이 실제로는 해를 끼치는 경우를 나타내는 쌍으로 연결된 변형들이 사용되었습니다. 세 가지 주요 결과가 도출되었습니다. 첫째, 취약한 안전성은 특정 상황에만 해당되는 문제이며, 12개의 모델 모두에서 안전성과 상식 간의 격차(평균 +17.4pp)가 관찰되었습니다. 기본적인 정확도는 취약성을 예측하지 못합니다. 90% 이상의 기본 정확도를 보이는 모델들 중에서도 취약성 발생률은 13.7%에서 90.0%까지 다양합니다. 둘째, 실패 원인은 오해 때문이 아니라 정책 재정의 때문입니다. 모델들은 모든 경우에 상황 변화를 인지하고 있음에도 불구하고, 업데이트 유형과 모델 패밀리에 따라 세 가지 서로 다른 메커니즘을 통해 지속적으로 잘못된 행동을 수행합니다. 셋째, 재앙적인 결과 변화 시나리오에 대한 수동 검토에서, 표준적인 액션 레벨 가이드라인은 단 하나의 경우도 잡아내지 못하는 반면, 상태 정보를 활용하는 검증기는 모든 경우를 정확하게 잡아내면서 오탐은 발생하지 않았습니다. 이는 액션 레벨의 콘텐츠 필터링이 결과 변화에 대해 체계적으로 맹목적이라는 것을 시사하며, 상태 정보를 고려하는 아키텍처 대안의 필요성을 강조합니다. 우리는 연구 프로토콜, 변형된 벤치마크 및 배포 검사를 공개합니다.

Original Abstract

Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational update flips which action is safe. We term this failure brittle safety. To diagnose it, we introduce context-flip evaluation, testing 12 models across a safety benchmark (PacifAIst) and two commonsense controls using paired variants where the nominally safe action produces harm. Three findings emerge. First, brittle safety is safety-specific: all 12 models exhibit a safety-commonsense gap (mean +17.4 pp). Baseline accuracy fails to predict brittleness: among models above 90% baseline accuracy, brittleness rates range from 13.7% to 90.0%. Second, failures stem from policy override rather than miscomprehension: despite acknowledging the context change in every case, models persist via three distinct mechanisms that vary by update type and model family. Third, on a hand-audited probe of catastrophic consequence-flip scenarios, standard action-level guardrails catch none, while a state-aware validator catches all without false alarms on correct interventions. This indicates that action-level content moderation is systematically blind to consequence-flips, motivating state-aware architectural alternatives. We release our protocol, perturbed benchmarks, and deployment probe.

3 Citations
0 Influential
0.5 Altmetric
5.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!