2605.29396v1 May 28, 2026 cs.AI

정렬되었지만 취약한: 제로차수 최적화를 통한 LLM 안전성 강건성 향상

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization

Jian Lou
Jian Lou
Citations: 21
h-index: 3
Yifan Wu
Yifan Wu
Citations: 38
h-index: 1
Dingzirui Wang
Dingzirui Wang
Citations: 251
h-index: 9
Yuxi Zhou
Yuxi Zhou
Citations: 38
h-index: 4
Zhihao Liu
Zhihao Liu
Citations: 47
h-index: 3
Yuke Hu
Yuke Hu
Citations: 222
h-index: 8

대규모 언어 모델(LLM)의 안전 정렬은 유해하거나 위험한 행동을 줄이는 동시에 일반적인 활용성을 유지하는 것을 목표로 합니다. 그러나 최근 연구 결과에 따르면, 정렬 효과는 취약할 수 있습니다. 파라미터 노이즈, 활성화 노이즈 또는 양자화와 같은 가벼운 사후 정렬 조작은 의도된 안전 행동을 쉽게 약화시킬 수 있습니다. 기존의 강건성 향상 노력은 주로 데이터 큐레이션, 수정된 정렬 목표 및 안전에 중요한 파라미터 식별에 집중되었으며, 최적화기의 역할 자체는 거의 탐구되지 않았습니다. 본 논문에서는 기본 최적화기 관점에서 안전 정렬의 강건성을 처음으로 연구합니다. 이러한 최적화기 중심적인 관점은 자연스럽게 제로차수 최적화를 지목하며, 이는 안전 정렬을 변화에 적용하여 강건성 관련 신호를 제공합니다. 이 통찰력을 바탕으로, 우리는 표준 1차 정렬을 먼저 수행하고 그 다음 제로차수 수정을 적용하여 강건성을 향상시키는 하이브리드 프레임워크를 제안합니다. 이론적으로나 실험적으로, 소수의 제로차수 수정 단계만으로도 강건성을 향상시키면서 안전 정렬을 유지할 수 있음을 보여줍니다. 또한, 우리는 제로차수 수정을 수행하는 과정에서 발생하는 내재적인 변화 기반 평가를 활용하여 계층별 강건성 민감도를 추정함으로써, 훈련 오버헤드를 최소화하면서 업데이트를 강건성이 중요한 계층에 집중하도록 하여 제로차수 수정의 효율성을 더욱 향상시킵니다.

Original Abstract

Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings reveal that alignment effects can be fragile: lightweight post-alignment manipulations, such as parameter noise, activation noise, or quantization, can easily weaken the intended safety behavior. Prior efforts to improve robustness have primarily focused on data curation, modified alignment objectives, and safety-critical parameter identification, leaving the role of the optimizer itself largely unexplored. In this paper, we are the first to study the robustness of safety alignment from the perspective of the base optimizer. This optimizer-centric view naturally points to zeroth-order optimization, which provides a robustness-oriented signal by evaluating safety alignment under perturbations. Based on this insight, we propose a hybrid framework that first performs standard first-order safety alignment and then applies zeroth-order refinement to improve robustness. Both theoretically and empirically, we show that only a few zeroth-order refinement steps can enhance robustness while preserving safety alignment. We further improve the efficiency of zeroth-order refinement by exploiting its inherent perturbation-based evaluations to estimate layer-wise robustness sensitivity, enabling the refinement process to concentrate updates on robustness-critical layers with modest training overhead.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!