2509.03985v2 Sep 04, 2025 cs.CR

NeuroBreak: 대규모 언어 모델의 내부 제약 우회 메커니즘 분석

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

Yingcai Wu
Yingcai Wu
Citations: 108
h-index: 3
Ye Zhang
Ye Zhang
Citations: 138
h-index: 3
Tianyu Du
Tianyu Du
Citations: 1,839
h-index: 17
Shouling Ji
Shouling Ji
Citations: 74
h-index: 5
Chuhan Zhang
Chuhan Zhang
Citations: 18
h-index: 2
Bowen Shi
Bowen Shi
Citations: 2
h-index: 1
Yuyou Gan
Yuyou Gan
Citations: 107
h-index: 3
Dazhang Deng
Dazhang Deng
Citations: 29
h-index: 3

대규모 언어 모델(LLM)은 일반적으로 배포 및 사용 시 불법적이거나 비윤리적인 출력을 방지하기 위한 안전성 정렬 과정을 거칩니다. 그러나, 적대적 프롬프트를 사용하여 안전 장치를 우회하도록 설계된 제약 우회 공격 기술이 지속적으로 발전하면서 LLM의 보안 방어에 대한 압박이 커지고 있습니다. 제약 우회 공격에 대한 저항력을 강화하려면 LLM의 보안 메커니즘과 취약점에 대한 심층적인 이해가 필요합니다. 그러나, LLM의 막대한 수의 매개변수와 복잡한 구조는 내부 관점에서 보안 취약점을 분석하는 것을 어렵게 만듭니다. 본 논문에서는 신경 수준의 안전 메커니즘을 분석하고 취약점을 완화하기 위해 설계된 상향식 제약 우회 분석 시스템인 NeuroBreak를 제시합니다. 저희는 AI 보안 분야의 세 명 전문가와의 협력을 통해 시스템 요구 사항을 신중하게 설계했습니다. 이 시스템은 다양한 제약 우회 공격 방법을 종합적으로 분석합니다. NeuroBreak는 계층별 표현 탐색 분석을 통합하여 모델이 생성 과정 전반에 걸쳐 의사 결정을 내리는 방식에 대한 새로운 관점을 제공합니다. 또한, 이 시스템은 의미론적 및 기능적 관점에서 중요한 뉴런을 분석하는 기능을 지원하며, 이를 통해 보안 메커니즘에 대한 더 깊은 탐색을 용이하게 합니다. 저희는 시스템의 효과를 검증하기 위해 정량적 평가와 사례 연구를 수행하고, 진화하는 제약 우회 공격에 대한 차세대 방어 전략 개발을 위한 기계적인 통찰력을 제공합니다.

Original Abstract

In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypass safety mechanisms with adversarial prompts, has placed increasing pressure on the security defenses of LLMs. Strengthening resistance to jailbreak attacks requires an in-depth understanding of the security mechanisms and vulnerabilities of LLMs. However, the vast number of parameters and complex structure of LLMs make analyzing security weaknesses from an internal perspective a challenging task. This paper presents NeuroBreak, a top-down jailbreak analysis system designed to analyze neuron-level safety mechanisms and mitigate vulnerabilities. We carefully design system requirements through collaboration with three experts in the field of AI security. The system provides a comprehensive analysis of various jailbreak attack methods. By incorporating layer-wise representation probing analysis, NeuroBreak offers a novel perspective on the model's decision-making process throughout its generation steps. Furthermore, the system supports the analysis of critical neurons from both semantic and functional perspectives, facilitating a deeper exploration of security mechanisms. We conduct quantitative evaluations and case studies to verify the effectiveness of our system, offering mechanistic insights for developing next-generation defense strategies against evolving jailbreak attacks.

2 Citations
0 Influential
8.5 Altmetric
44.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!