LLM의 작동 원리 분석: 적대적 환경에서 레이어별 특징 취약점 연구
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
안전 정렬 과정을 거치더라도, 대규모 언어 모델(LLM)은 여전히 악의적인 결과를 생성하도록 조작될 수 있습니다. 기존 공격 연구는 이러한 취약점을 보여주었지만, 그 원인이 되는 내부 메커니즘은 밝혀지지 않았습니다. 본 연구는 제로샷 공격 성공 여부가 프롬프트뿐만 아니라 특정 내부 특징에 의해 결정되는지 질문합니다. 곰마-2-2B 모델과 BeaverTails 데이터셋을 사용하여 세 단계로 구성된 파이프라인을 제안합니다. 첫째, 부분 공간 유사성을 이용하여 적대적 응답에서 개념과 관련된 토큰을 추출합니다. 둘째, 세 가지 특징 그룹화 전략(클러스터링, 계층적 연결, 단일 토큰 기반)을 적용하여 모든 26개 모델 레이어에서 관련된 토큰의 SAE(Self-Attention Embedding) 특징 하위 그룹을 식별합니다. 셋째, 식별된 각 하위 그룹에서 가장 중요한 특징을 증폭시켜 모델의 작동을 조작하고, 표준화된 LLM 판별 시스템을 사용하여 유해성 점수의 변화를 측정합니다. 세 가지 접근 방식 모두에서, [16-25] 레이어의 특징이 조작에 상대적으로 더 취약했습니다. 세 가지 방법 모두 중간에서 후반 레이어의 특징 하위 그룹이 안전하지 않은 결과 생성에 더 큰 영향을 미친다는 것을 확인했습니다. 이러한 결과는 곰마-2-2B 모델의 제로샷 공격 취약점이 중간에서 후반 레이어의 특징 하위 그룹에 국한되어 있다는 증거를 제공하며, 이는 현재 프롬프트 수준의 방어보다 특징 수준의 표적 개입이 적대적 견고성을 향상시키는 더욱 효과적인 방법이 될 수 있음을 시사합니다.
Large language models (LLMs) can still be jailbroken into producing harmful outputs despite safety alignment. Existing attacks show this vulnerability, but not the internal mechanisms that cause it. This study asks whether jailbreak success is driven by identifiable internal features rather than prompts alone. We propose a three-stage pipeline for Gemma-2-2B using the BeaverTails dataset. First, we extract concept-aligned tokens from adversarial responses via subspace similarity. Second, we apply three feature-grouping strategies (cluster, hierarchical-linkage, and single-token-driven) to identify SAE feature subgroups for the aligned tokens across all 26 model layers. Third, we steer the model by amplifying the top features from each identified subgroup and measure the change in harmfulness score using a standardized LLM-judge scoring protocol. In all three approaches, the features in the layers [16-25] were relatively more vulnerable to steering. All three methods confirmed that mid to later layer feature subgroups are more responsible for unsafe outputs. These results provide evidence that the jailbreak vulnerability in Gemma-2-2B is localized to feature subgroups of mid to later layers, suggesting that targeted feature-level interventions may offer a more principled path to adversarial robustness than current prompt-level defenses.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.