2606.05614v1 Jun 04, 2026 cs.AI

안전 역설: 강화된 안전 인식의 부작용 - LLM이 후처리 공격에 취약해지는 이유

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

Shaoyang Xu
Shaoyang Xu
Citations: 186
h-index: 8
Wenxuan Zhang
Wenxuan Zhang
Citations: 9
h-index: 1
Long Hoang
Long Hoang
Citations: 60
h-index: 3
H. V. Le
H. V. Le
Citations: 44
h-index: 5
Wei Lu
Wei Lu
Citations: 882
h-index: 5

대규모 언어 모델(LLM)은 유해한 요청을 거부하도록 엄격하게 조정되며, 이 과정에서 잠재적으로 위험한 콘텐츠를 평가하고 식별하는 능력이 발달합니다. 본 연구에서는 이러한 고도화된 안전 인식이 의도치 않게 치명적인 취약점을 야기한다는 것을 밝힙니다. 우리는 단일 질의로 LLM의 방어 체계를 우회하는 '후처리 공격(Posterior Attack)'을 제안합니다. 이 공격은 모델이 내부 분류기가 일반적으로 유해하다고 판단할 정확한 응답을 생성하도록 유도합니다. 30개의 공개 소스 LLM(최대 350억 파라미터) 및 최첨단 모델(예: GPT-5, Claude 4.6)에 대한 광범위한 실험적 평가를 통해 놀라운 현상을 관찰했습니다. 즉, 안전 판단 능력이 뛰어난 모델일수록 이 공격에 더 취약합니다. 이러한 현상을 설명하기 위해 우리는 '안전 역설'을 정립하고, 안전 조정의 단조로운 개선이 자연스럽게 후처리 취약성을 증폭시킨다는 것을 분석적으로 보여줍니다. 또한 강화 학습 개입을 통해 인과 관계를 확립했으며, 모델의 안전 판단 능력을 의도적으로 저하시키면 공격에 대한 면역력이 생기는 반면, 판단 능력을 향상시키면 취약성이 더욱 악화된다는 것을 입증했습니다. 우리의 연구 결과는 현재의 조정 패러다임에 잠재적인 결함이 있음을 시사하며, 방어 메커니즘이 추가적인 구조적 개선이 필요할 수 있습니다.

Original Abstract

Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced safety awareness inadvertently introduces a fatal vulnerability. We introduce Posterior Attack, a single-query jailbreak that bypasses guardrails by prompting the model to generate the exact harmful response its internal classifier would normally flag as unsafe. Through extensive empirical evaluation across 30 open-source LLMs (up to 35B parameters in size) and frontier models (e.g., GPT-5, Claude 4.6), we observe a striking phenomenon: models with superior safety-judgment capabilities are disproportionately more susceptible to this exploitation. To explain this, we formalize the Safety Paradox, analytically showing that monotonic improvements in safety alignment naturally amplify posterior vulnerability. Finally, we establish a causal link via reinforcement learning interventions, exemplifying that artificially degrading a model's safety judgment immunizes it against the attack, whereas enhancing judgment exacerbates the vulnerability. Our findings highlight potential flaws in current alignment paradigms, indicating that defense mechanisms may require further structural refinement.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!