2606.16751v1 Jun 15, 2026 cs.CR

다중 방어 전략을 겨냥한 자동화된 제이크 공격

Automated jailbreak attack targeting multiple defense strategies

Yanqing Li
Yanqing Li
Citations: 0
h-index: 0
Qi Wang
Qi Wang
Citations: 13
h-index: 2
Chengcheng Wan
Chengcheng Wan
Citations: 265
h-index: 8
Weijia He
Weijia He
Citations: 490
h-index: 7
Hanqi Sun
Hanqi Sun
Citations: 5
h-index: 1
Xiaodong Gu
Xiaodong Gu
Citations: 302
h-index: 9
Jiangtao Wang
Jiangtao Wang
Citations: 0
h-index: 0

대규모 언어 모델(LLM)은 다양한 작업에서 놀라운 능력을 보여주었지만, 적대적 프롬프트 기반 공격에 취약하기 때문에 안전성 문제가 중요한 과제로 남아 있습니다. 본 논문에서는 방어를 고려한 관점에서 효과적인 블랙박스 공격 프롬프트를 체계적으로 구성하도록 설계된 적대적 테스트 프레임워크인 UNIATTACK을 제시합니다. 기존 접근 방식이 정적 템플릿이나 반복적인 모델별 튜닝에 의존하는 것과는 달리, UNIATTACK은 다양한 기존 공격에서 최소한의 영향력을 갖는 공격 특징을 추출하고, 특수 목적의 공격 LLM을 통해 이를 최적화하며, 자동화된 개선 과정을 통해 유연한 템플릿으로 구성합니다. 이러한 특징 중심적인 구성 방식을 통해, UNIATTACK은 여러 모델과 안전 범주에 걸쳐 일반화되는 단일 시도 공격을 가능하게 하며, LLM의 견고성을 평가하는 실용적인 도구를 제공합니다. 실험 결과, UNIATTACK은 다층 방어 메커니즘이 적용된 모델에서 기준 방법보다 평균 공격 성공률(ASR)을 64.63%~248.82% 향상시켰으며, 비용은 기준 방법의 0.03%~4.96%에 불과합니다. UNIATTACK 관련 자료는 https://anonymous.4open.science/r/UniAttack-Artifact-30F1 에서 확인할 수 있습니다.

Original Abstract

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their safety remains a critical concern due to their susceptibility to adversarial prompt-based attacks. In this paper, we present UNIATTACK, an adversarial testing framework designed from a defense-oriented perspective to systematically construct effective black-box attack prompts. Unlike prior approaches that rely on static templates or iterative model-specific tuning, UNIATTACK extracts minimal but high-impact attack features from diverse existing attacks, optimizes them via a specialized attacker LLM, and composes them into flexible templates through automated refinement process. This feature-centric construction enables one-shot attacks that generalize across multiple models and safety categories, providing a practical tool for assessing LLM robustness. Our evaluation results shows that compared to the baselines, UNIATTACK achieves an average attack success rate (ASR) improvement of 64.63\%-248.82\% on models deployed with multi-layered defense mechanisms and it only takes 0.03\%-4.96\% cost of the baselines. UNIATTACK artifact is available at https://anonymous.4open.science/r/UniAttack-Artifact-30F1.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!