감정이 중요하다: 구문 민감성이 안전 정렬을 저해하는 방식
Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment
대규모 언어 모델은 일반적으로 안전 정책에 맞추기 위해 추가 학습을 거치지만, 기존의 안전 장치를 우회하는 다양한 정교한 공격 기법들이 존재합니다. 예를 들어, Andriushchenko et al. (2025)의 이전 연구에서는 현재 시제에서 과거 시제로 문장 시제를 변경하는 것만으로도 유해한 응답을 이끌어낼 수 있다는 사실이 밝혀졌습니다. 본 연구에서는 명령형이 아닌 더 일반적인 구문 형태에서의 취약점을 발견했습니다. 행동 평가를 통해 최대 700억 개의 파라미터를 가진 16개의 모델에서 이러한 구문적 취약점이 존재함을 확인했습니다. 근본 원인을 조사하기 위해 인과 매개 분석을 수행한 결과, 거부 결정은 부분적으로 상위 수준의 구문 특징에 의해 영향을 받는다는 것을 발견했습니다. 이러한 순수한 구문 특징을 조작하여 거부를 유발하고 억제할 수 있었습니다. 마지막으로, 우리는 이러한 부적절한 작동 방식이 오픈 소스 모델의 언어 편향적인 추가 학습 데이터에서 비롯되었음을 추적했으며, 구문 다양성을 증가시키면 이 문제를 완화할 수 있음을 보여주었습니다. 우리의 연구 결과는 현재의 정렬 접근 방식이 거부 결정의 순수한 의미론적 기반을 방해하는 혼란 요소를 도입한다는 것을 시사합니다.
Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70B parameters, using behavioral evaluation. To investigate the root cause, we apply causal mediation analysis, finding that refusal is partially conditioned on upstream syntactic features. By steering these purely syntactic features we are able to trigger and suppress refusal. Finally, we trace this ill-conditioning to linguistically biased post-training data of open-source models and show that increasing syntactic diversity can mitigate the issue. Our findings suggest that current alignment approaches introduce confounders that prevent a pure semantic grounding of the refusal decision.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.