2605.05995v1 May 07, 2026 cs.CR

안전 앵커: 기하학적 병목 현상을 활용한 유해한 미세 조정 방어

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

Huaiyu Dai
Huaiyu Dai
Citations: 16
h-index: 2
Guoxin Lu
Guoxin Lu
Citations: 3
h-index: 1
Letian Sha
Letian Sha
Citations: 30
h-index: 3
Peijie Sun
Peijie Sun
Nanjing University of Posts and Telecommunications
Citations: 2,156
h-index: 20
Fu Xiao
Fu Xiao
Citations: 5
h-index: 1
Qing Wang
Qing Wang
Citations: 10
h-index: 2
Hao Zhou
Hao Zhou
Citations: 74
h-index: 3

대규모 언어 모델(LLM)의 안전 정렬은 여전히 유해한 미세 조정(HFT)에 취약합니다. 기존의 방어 기법들은 파라미터, 기울기 또는 내부 표현에 제약을 가하지만, 지속적인 HFT 공격에 의해 효과적으로 무력화될 수 있습니다. 우리의 분석에 따르면, 이는 고차원 파라미터 공간의 내재된 중복성 때문입니다. 공격자들은 방어 제약 조건에 수직인 최적화 경로를 활용하여 유해한 기능을 복원하면서도 안전 제한을 준수하는 것처럼 보이게 만듭니다. 이러한 문제를 해결하기 위해, 우리는 안전 병목 현상 정규화(SBR)를 제안합니다. SBR은 방어적 초점을 중복된 파라미터 공간에서 벗어나, 기하학적 병목 현상 역할을 하는 언임베딩 레이어로 전환합니다. SBR은 유해한 쿼리의 최종 은닉 상태를 안전하게 정렬된 모델의 상태에 고정함으로써, 모델이 지속적인 HFT 공격 하에서도 안전한 응답을 유지할 수 있도록 합니다. 광범위한 실험 결과는 SBR의 효과를 확인했으며, 단 하나의 안전 앵커를 사용하는 것만으로도 유해성 점수를 10 이하로 줄이면서도 안전한 작업에서의 경쟁력 있는 성능을 유지할 수 있음을 보여줍니다.

Original Abstract

The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constraints on parameters, gradients, or internal representations, we observe that they can be effectively circumvented under persistent HFT. Our analysis traces this failure to the inherent redundancy of the high-dimensional parameter space: attackers exploit optimization trajectories that are orthogonal to defense constraints to restore harmful capabilities while deceptively adhering to safety restrictions. To address this, we propose Safety Bottleneck Regularization (SBR). SBR shifts the defensive focus from the redundant parameter space to the unembedding layer, which serves as a geometric bottleneck. By anchoring the final hidden states of harmful queries to those of the safety-aligned model, SBR enables the model to maintain safe responses even under persistent HFT. Extensive experiments confirm SBR's effectiveness, demonstrating that utilizing just a single safety anchor is sufficient to reduce the Harmful Score to $<$10 while preserving competitive performance on benign downstream tasks.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!