2608.08542v1 Aug 09, 2026 cs.LG

기술과 안전의 만남: 기술 통합 언어 모델의 적응형 탈옥 방어 능력에 대한 벤치마킹 및 특성 분석

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

Xinran Xu
Xinran Xu
Citations: 17
h-index: 2
Yu Ma
Yu Ma
Citations: 12
h-index: 2
Hongli Shi
Hongli Shi
Citations: 0
h-index: 0
Jing Li
Jing Li
Citations: 0
h-index: 0
Weiwei Hou
Weiwei Hou
Citations: 0
h-index: 0

모델 병합은 기존 언어 모델을 재학습하지 않고 새로운 기능을 부여하는 가장 일반적인 방법이 되었습니다. 실무자는 수학, 코딩 또는 전문 분야의 작업 벡터를 작업 산술, TIES 또는 DARE 기술을 사용하여 안전하게 조정된 기본 모델에 통합합니다. 이러한 편리함은 안전상의 위험을 초래할 수 있지만, 대부분의 증거는 정적 거부 테스트에 기반하며, 이는 미리 정의된 유해한 프롬프트를 사용하여 준수 여부를 평가하는 방식입니다. 우리는 이러한 접근법이 오해를 불러일으킬 수 있다고 주장합니다. 안전 조정은 주로 생성되는 처음 몇 개의 토큰에 집중되어 있기 때문에, 병합된 모델의 정적 거부 성능이 양호하더라도 실제 적응 공격으로 인해 쉽게 탈옥될 수 있습니다. 우리는 SkillSafe-Bench라는 제어된 벤치마크를 소개하며, 이 벤치마크는 기술 통합 모델을 정적 거부 능력, 적응형 탈옥 방어 능력 및 보수적인 두 명의 심사위원의 AND 규칙 하에서의 기능 유지 능력에 대해 평가합니다. 여섯 가지 공개 웨이트 기반 모델(다섯 개의 제품군, 두 가지 규모)을 대상으로 실험한 결과, 정적 안전성은 공격에 대한 견고성을 예측하지 못했습니다. 의미론적 템플릿 공격 시, 취약한 기본 모델(Qwen의 모든 규모 및 Gemma)에서 '안전해 보이는' 병합 모델은 60-76%의 탈옥 성공률을 보인 반면, 다른 모델(Llama, Phi-4)은 견고성을 유지했습니다. 또한, 병합의 정적 효과는 기본 모델에 따라 달라지며, 데이터 없이 얻은 기하학적 신호(작업 벡터와 안전 공간 간의 중첩 영역)를 통해 동일한 레시피로 인해 발생하는 안전성 저하 현상을 분석하고, 기능을 유지하면서 이러한 저하를 제거하는 SubSafe-Merge 기술을 제안합니다. 병합된 언어 모델에 대한 적응형 평가는 선택 사항이 아니며, 정적 검사에서 안전해 보이는 모델일수록 더욱 필요합니다.

Original Abstract

Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE. This convenience is known to carry a safety cost, but almost all of that evidence rests on static refusal tests: fixed harmful prompts scored for compliance. We argue this is misleading. Because safety alignment is "shallow," concentrated in the first few generated tokens, a merged model's static refusal can stay clean while a real adaptive attack still breaks it. We introduce SkillSafe-Bench, a controlled benchmark that scores skill-merged models on static refusal, adaptive jailbreak robustness, and capability retention under a conservative two-judge AND rule. Across six open-weight bases (five families, two scales), static safety does not predict robustness to attack: under a semantic template attack, safe-looking merges on the fragile bases (both Qwen scales and Gemma) are jailbroken 60-76% of the time while others (Llama, Phi-4) stay robust. We further show the static effect of merging is base-conditional, characterize same-recipe abliteration-style safety erosion through a data-free geometric signal (the overlap of a task vector with a safety subspace), and outline SubSafe-Merge, which projects this overlap away to remove that erosion at held capability. Adaptive evaluation is not optional for merged LLMs: the models that most need it look safe under static screening.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!