2608.03210v1 Aug 04, 2026 cs.CL

ICO: 반복적인 컨텍스트 최적화를 통한 의미 변화 기반의 안전 우회 공격 성능 향상

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

Simeng Qin
Simeng Qin
Citations: 132
h-index: 7
Qing Guo
Qing Guo
Citations: 184
h-index: 9
G. Pu
G. Pu
Citations: 3,531
h-index: 30
Xinfeng Li
Xinfeng Li
Citations: 1,241
h-index: 18
Hujian Zhu
Hujian Zhu
Citations: 0
h-index: 0
Felix Juefei-Xu
Felix Juefei-Xu
Citations: 1,068
h-index: 18
Peng Zeng
Peng Zeng
Citations: 9
h-index: 2
Yihao Huang
Yihao Huang
Citations: 1,097
h-index: 19

기초 모델은 다양한 작업에서 놀라운 성공을 거두었지만, 여전히 취약점을 가지고 있습니다. 이러한 취약점을 조사하기 위해 최근에는 의미 변화를 이용한 안전 우회 공격이 유망한 공격 패러다임으로 등장했습니다. 이 방법은 원래의 유해한 질문에서 유해한 용어를 양 benign적인 대안으로 대체하고, 컨텍스트 정보를 활용하여 대상 모델이 이러한 대안을 해당 유해한 개념으로 재해석하도록 유도하여 명시적인 안전 장치를 우회합니다. 그러나 기존의 의미 변화 기반 안전 우회 공격은 종종 제한된 효과를 보입니다. 본 연구에서는 이러한 한계가 컨텍스트의 의미 변화 능력을 간과했기 때문에 발생한다는 것을 밝힙니다. 체계적인 분석을 통해, 컨텍스트는 상당히 다른 수준의 의미 변화 유도 능력을 가지고 있으며, 더 강력한 의미 변화 능력을 가진 컨텍스트일수록 모델이 유해한 의미를 복원하고 안전 우회 공격에 성공할 가능성이 높다는 것을 확인했습니다. 이러한 발견을 바탕으로, 우리는 효과적인 컨텍스트의 특징을 체계적으로 식별하고 추출했으며, 반복적인 컨텍스트 최적화(ICO) 기능을 갖춘 블랙박스 기반의 컨텍스트 인식 의미 변화 안전 우회 프레임워크를 제안합니다. ICO는 각 단계에서 이러한 특징과 대상 모델로부터 얻은 피드백을 활용하여 컨텍스트를 최적화합니다. 세 가지 데이터셋과 여덟 개의 기초 모델에 대한 광범위한 실험 결과, ICO가 8개의 최첨단 기준 모델보다 일관되게 우수한 성능을 보였으며, 평균 공격 성공률이 74.6%에 달했습니다.

Original Abstract

Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass explicit safety mechanisms by replacing harmful terms in original harmful questions with benign alternatives and leveraging contextual information to induce the target model to reinterpret these alternatives as their corresponding harmful concepts. However, existing semantic-shift jailbreaks often achieve limited effectiveness. In this work, we reveal that this limitation arises from overlooking the semantic-shift capability of contexts. Through systematic analysis, we find that contexts exhibit substantially different abilities in inducing semantic shifts: contexts with stronger semantic-shift capabilities are more likely to guide models toward recovering harmful meanings and achieving successful jailbreaks. Based on this finding, we systematically identify and distill the characteristics of effective contexts and propose a black-box context-aware semantic-shift jailbreak framework with Iterative Context Optimization (ICO). In each iteration, ICO leverages these characteristics and feedback from the target model to optimize contexts. Extensive experiments on three datasets and eight target foundation models demonstrate that ICO consistently outperforms eight state-of-the-art baselines, achieving an average attack success rate of 74.6%.

0 Citations
0 Influential
15 Altmetric
75.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!