SafeNexus: 멀티모달 대규모 언어 모델(MLLM)에서 모드에 관계없이 안전성을 확보하는 뉴런 발견 및 제어
SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs
대규모 언어 모델(LLM)은 유망한 안전 성능을 보여주었지만, 이를 멀티모달 대규모 언어 모델(MLLM)로 확장하면 향상된 멀티모달 기능과 기존의 안전 메커니즘 간에 상당한 격차가 발생합니다. 현재의 방어 기술은 대부분 특정 모드 설정에 국한되어 있어, 더 광범위한 교차 모드 위협에 대한 견고성이 제한됩니다. 이러한 격차를 해소하기 위해, 뉴런 수준의 개입 전략을 채택하는 교차 모드 안전 정렬 프레임워크인 SafeNexus를 제안합니다. 먼저, 중간 레이어 활성화 패턴을 분석하고 중요도 점수를 통해 기능적 중요성을 정량화하여 기능적으로 특수화된 뉴런을 식별하는 뉴런 위치 추론 패러다임을 정의합니다. 이 패러다임을 바탕으로, 대비 데이터를 활용하여 모드에 종속적인 안전 뉴런(BS-뉴런)을 식별하고, 표적 억제를 통해 각 모드 내에서 안전 행동을 조절하는 역할을 검증합니다. 추가적으로 교차 모드 분석을 통해 개별 모드에서 식별된 BS-뉴런의 공통 부분인 모드 보편 안전 뉴런(US-뉴런)을 정의하며, 이는 유해한 교차 모드 공격에 대한 방어의 핵심으로 사용됩니다. 저희는 이러한 뉴런을 억제하면 여러 모드에서 안전 성능이 크게 저하되지만, 전체적인 유용성은 비교적 큰 영향을 받지 않는다는 것을 확인했습니다. 이러한 통찰력을 바탕으로, 활성화 수준 안전 증폭기 및 안전 뉴런 조정기라는 두 가지 안전 정렬 전략을 제안합니다. 제안된 전략은 모델의 안전성을 향상시키는 두 가지 뚜렷한 경로를 제공합니다. 첫 번째는 US-뉴런의 활성화 크기를 증폭시키고, 두 번째는 표적 미세 조정을 통해 선택적으로 조정합니다. 광범위한 실험 결과, 저희 방법이 다양한 모드 조합을 포괄하는 안전 벤치마크에서 기존 최고 성능 기술보다 우수한 성능을 보이며, 동시에 유용성을 효과적으로 유지한다는 것을 입증했습니다.
Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes a significant gap between expanded multimodal capabilities and existing safety mechanisms. Current defenses remain predominantly confined to specific modal settings, thereby limiting their robustness against broader cross-modal threats. To bridge this gap, we introduce SafeNexus, a cross-modal safety alignment framework that adopts a dedicated neuron-level intervention strategy. First, we formulate a neuron localization paradigm that identifies functionally specialized neurons by characterizing intermediate-layer activation patterns and quantifying their functional salience through importance scoring. Building upon this paradigm, we exploit contrastive data to identify modality-bound safety neurons (BS-Neurons), and validate their role in regulating safety behavior within each modality via targeted suppression. Further cross-modal analysis defines modality-universal safety neurons (US-Neurons) as the shared subset of BS-Neurons identified across individual modalities, serving as the core for defending against harmful cross-modal attacks. We observe that suppressing these neurons substantially degrades safety performance across modalities, while leaving overall utility largely unaffected. Building on these insights, we propose two safety alignment strategies: activation-level safety amplifier and safety neuron calibrator. The proposed strategies enhance model safety through two distinct routes: the former amplifies the activation magnitudes of US-Neurons, while the latter selectively calibrates them via targeted fine-tuning. Extensive experiments demonstrate that our method outperforms prevailing state-of-the-art approaches on safety benchmarks spanning diverse modality combinations, while effectively preserving utility.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.