거절이 안전해 보일 때: 안전 방어 모델에서의 거절 신호 단축키
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
안전 방어 시스템은 유해 콘텐츠를 필터링하는 데 널리 사용되며, 일반적으로 레이블이 지정된 프롬프트-응답 쌍을 사용하여 지도 학습 방식으로 훈련됩니다. 본 연구에서는 두 가지 널리 사용되는 안전 방어 훈련 데이터 세트인 WildGuardMix와 GR-Train을 분석한 결과, 유해한 프롬프트에 대한 응답 중에서 거절 표현은 거의 예외 없이 무해한 레이블과 함께 나타나는 경향이 있음을 확인했습니다. 이러한 불균형은 '거절 신호 단축키'라는 현상을 야기합니다. 즉, 유해한 응답에 거절 신호를 삽입하면 안전 방어 시스템의 판단이 유해에서 무해로 바뀔 수 있습니다. 이 단축키는 분석 대상 데이터 세트로 훈련된 모델뿐만 아니라 LlamaGuard3 및 Qwen3Guard와 같이 공식적으로 출시되었지만 훈련 데이터가 공개되지 않은 모델에도 영향을 미칩니다. 이러한 현상은 응답 위치에 관계없이 지속되며, 일반적으로 동일 계열 내의 작은 변형에서 더 강하게 나타납니다. 이러한 문제를 완화하기 위해, 본 연구에서는 재훈련 없이 작동하는 가벼운 사후 개입 방식으로 '희소 보완 마스킹'을 적용하여 단축키와 관련된 일부 어텐션 헤드 및 MLP 뉴런을 식별하고 억제합니다. 두 가지 주요 벤치마크에서 이러한 개입은 거절 신호로 인해 발생하는 응답 초반 탐지 실패를 약 79% 상대적으로 감소시키는 효과를 보였으며, 동시에 표준 탐지 성능은 유지되었습니다. 단일 응답 위치의 신호를 사용하여 최적화되었지만, 억제 효과는 새로운 위치 및 데이터 세트로 전달되는 경향이 있으며, 이는 응답 위치 간의 단축키 표현이 부분적으로 공유된 내부 구성 요소에 의해 매개됨을 시사합니다. 추가 분석 결과, 단축키 의존성과 정당한 거절 인식이 부분적으로 기능적으로 분리될 수 있다는 증거가 나타났습니다. 즉, 단축키를 억제하면 안전 방어 시스템의 진정한 거절 인식을 유지하는 데 도움이 됩니다.
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit two widely used safety-guard training datasets, WildGuardMix and GR-Train, and find that among responses to harmful prompts, refusal expressions co-occur almost exclusively with unharmful labels. This imbalance motivates what we term the refusal-cue shortcut: inserting a refusal cue into a harmful response could flip the guard's verdict from harmful to unharmful. The shortcut affects not only guards trained on these datasets but also officially released models such as LlamaGuard3 and Qwen3Guard whose training data is undisclosed. It persists across response positions and is generally stronger in smaller variants within a family. To mitigate it, we adapt sparse complementary masking as a lightweight post-hoc intervention that identifies and suppresses a small set of shortcut-associated attention heads and MLP neurons without retraining. On two primary benchmarks, the intervention achieves an approximately 79% relative reduction in response-initial detection failures induced by refusal cues, while preserving standard detection performance. Although optimized using cues at a single response position, the suppression effect transfers to unseen positions and datasets, suggesting that shortcut manifestations across positions are partly mediated by shared internal components. Further analysis provides evidence that shortcut reliance and legitimate refusal recognition are partially functionally separable, as suppressing the shortcut broadly preserves the guard's ability to recognize genuine refusals.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.