2605.27823v1 May 27, 2026 cs.CR

적대적인 프롬프트 분리: 강력한 LLM 보안을 위한 의미 그래프 기반 방어

Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security

Xiang Fang
Xiang Fang
Citations: 255
h-index: 11
Wanlong Fang
Wanlong Fang
Citations: 380
h-index: 14

최근 대규모 언어 모델(LLM)은 안전 장치를 우회하고 유해하거나 부적절한 출력을 생성하는 방식으로 의미적 모호성을 악용한 적대적인 프롬프트에 점점 더 취약해지고 있습니다. 이러한 공격, 즉 탈옥 및 프롬프트 주입은 보안이 중요한 응용 분야에서 LLM의 무결성과 가용성에 상당한 위험을 초래합니다. 본 논문에서는 Adversarial Prompt Disentanglement (APD)라는 새로운 방어 메커니즘을 제안합니다. APD는 LLM에 의해 처리되기 전에 입력 프롬프트 내의 악성 구성 요소를 사전에 식별하고 무력화합니다. APD 프레임워크는 세 가지 핵심 혁신을 통합합니다: (1) 상호 정보 기반 의미 분해 방법을 사용하여 적대적이고 안전한 프롬프트 구성 요소를 분리하여 통계적 독립성을 보장합니다; (2) 그래프 기반 의도 분류 접근 방식을 사용하여 스펙트럼 분석을 활용하여 프롬프트의 의미론적 패턴에서 악성 패턴을 감지합니다; (3) 실제 독성 및 탈옥 프롬프트 데이터 세트를 기반으로 훈련된 경량 트랜스포머 기반 분류기를 사용하여 효율적이고 정확한 적대적 의도 탐지를 가능하게 합니다. 다양한 적대적인 프롬프트를 포함하는 다양한 데이터 세트에서 평가 결과, APD는 우수한 견고성을 보여주며 유해한 출력 생성을 85% 이상 줄이는 동시에 모델 성능에 미치는 영향은 거의 없습니다. 이 프레임워크의 계산 효율성은 실시간 배포를 지원하여 LLM을 보호하는 데 실용적인 솔루션을 제공합니다. 본 연구는 새로운 공격 및 ML 시스템의 무결성에 대한 방법을 포함한 머신 러닝 보안의 중요한 과제를 해결하고, 프롬프트 기반 적대적 위협에 대한 확장 가능하고 윤리적으로 책임 있는 방어를 제시합니다.

Original Abstract

Large Language Models (LLMs) are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulting in harmful or inappropriate outputs. Such attacks, including jailbreaking and prompt injection, pose significant risks to the integrity and availability of LLMs in security-critical applications. This paper proposes the Adversarial Prompt Disentanglement (APD) framework, a novel defense mechanism that proactively identifies and neutralizes malicious components in input prompts before they are processed by the LLM. The APD framework integrates three key innovations: (1) a mutual information-based semantic decomposition method to isolate adversarial and benign prompt components, ensuring statistical independence; (2) a graph-based intent classification approach that leverages spectral analysis to detect malicious patterns in prompt semantics; and (3) a lightweight transformer-based classifier trained on real-world datasets of toxic and jailbreaking prompts, enabling efficient and accurate adversarial intent detection. Evaluated on diverse datasets containing adversarial prompts, APD demonstrates superior robustness, reducing harmful output generation by over 85\% while maintaining negligible impact on model performance. The framework's computational efficiency supports real-time deployment, making it a practical solution for securing LLMs. Our work addresses critical challenges in machine learning security on novel attacks and integrity methods for ML systems, and offers a scalable, ethically grounded defense against prompt-based adversarial threats.

26 Citations
0 Influential
7 Altmetric
61.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!