2607.24392v1 Jul 27, 2026 cs.CR

LLM 방어 기법이 부작용을 일으킬 때: 안전성, 성능 및 비용 간의 균형 분석

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Yun Peng
Yun Peng
Citations: 78
h-index: 3
Simin Chen
Simin Chen
Citations: 111
h-index: 5
Tong Zhang
Tong Zhang
Citations: 0
h-index: 0
Zexin Li
Zexin Li
Citations: 37
h-index: 2

대규모 언어 모델(LLM)을 보호하는 데 필수적인 제재 방어 기법은 모델의 유용성을 저하시키는 2차적인 비용을 초래할 수 있습니다. 본 연구에서는 이러한 방어 기법들의 균형을 성능 영향, 양성 입력에 대한 과도한 거부 반응, 그리고 추론 비용이라는 세 가지 측면에서 체계적으로 분석합니다. 우리는 방어 기법들을 단일 범주로 취급하는 대신, 작동 전략에 따라 분류하고, 이러한 전략들이 서로 다른 부작용 프로필과 어떻게 연관되는지 살펴봅니다. 최첨단 방어 방법, 널리 사용되는 벤치마크 데이터셋, 그리고 대표적인 오픈 소스 LLM을 사용하여 분석한 결과, 방어 기법은 거의 예외 없이 다운스트림 성능을 향상시키지 못하며, 대신 안전성 향상 효과를 가용성과 효율성을 희생하는 방식으로 나타냅니다. 특히, 규칙 기반 방어는 작업 성능을 가장 잘 유지하는 반면, 지나치게 보수적인 자기 성찰 방어는 과도한 거부 반응을 자주 유발하며, 다중 라운드 방어는 가장 큰 런타임 오버헤드를 발생시키는 것으로 나타났습니다. 이러한 결과는 방어 기법의 부작용을 평가하기 위한 기준으로 활용될 수 있으며, 실제 배포 환경에서의 방어 기법 선택에 대한 실질적인 지침을 제공합니다.

Original Abstract

Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost. Rather than treating defenses as a single class, we organize them by operational strategy and examine how different strategies correlate with different side-effect profiles. Across state-of-the-art defense methods, widely used benchmark datasets, and representative open-source LLMs, we find that defenses rarely improve downstream capability, but instead vary in how they trade safety gains against usability and efficiency. In particular, rule-based defenses best preserve task performance, highly conservative self-reflective defenses often increase over-refusal, and multi-round defenses incur the largest runtime overhead. These results provide both a benchmark for evaluating defense side effects and practical guidance for selecting defenses under deployment constraints.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!