LLM 안전을 위한 온-폴리시 증류: 템플릿 불일치에 강건한 재정렬을 위한 라우팅 접근 방식
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
파인튜닝은 대규모 언어 모델(LLM)을 특화하는 주요 방법이지만, 중요한 취약점을 드러냅니다. 악의적인 데이터 제공자는 유해한 행동을 다운스트림 데이터에 삽입하여 전문적인 기술을 유지하면서도 필요에 따라 인간의 가치를 위반하는 모델을 만들 수 있습니다. 기존 안전 재정렬 방어 기법은 종종 세 가지 주요 한계로 인해 실제에서 효과가 미흡합니다. 첫째, 특수화된 기술의 파국적인 손실을 초래하는 경우가 많습니다. 둘째, 방어자가 공격자의 프롬프트 템플릿을 관찰할 수 없을 때 효과가 급격히 감소합니다. 셋째, 성공적으로 재정렬된 모델도 간단한 시스템 프롬프트 변경을 통해 쉽게 재탈옥될 수 있습니다. 이러한 문제점을 해결하기 위해, 우리는 정렬된 출력 확률 분포와 손상된 출력 확률 분포 간의 차이를 모델링하는 새로운 재정렬 프레임워크인 라우팅 기반 온-폴리시 증류(Routing-based On-Policy Distillation, ROPD)를 제안합니다. 다양한 수준의 정렬 강도를 가진 세 가지 기본 모델과 세 개의 데이터셋을 사용하여 ROPD를 최첨단 방어 기법 4가지와 비교하는 광범위한 실험을 수행했습니다. 우리의 결과는 기존 방어 기법이 템플릿 불일치를 마주하면, 다운스트림 작업 성능 저하가 심각하게 발생하는 것을 보여줍니다. 반대로, ROPD는 템플릿 불일치 위험을 크게 완화하고, 방어 효과와 기능 유지 모두에서 뛰어난 강건성을 유지합니다. 우리의 분석에 따르면 ROPD는 템플릿 변화에 완전히 면역적이지 않지만, 기존 방법과 비교했을 때 성능 저하가 미미하며, 이는 견고한 LLM 재정렬을 위한 새로운 기준을 제시합니다.
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender cannot observe the attacker's prompt template; and successfully realigned models remain susceptible to re-jailbreaking via simple system prompt switches. To address these challenges, we propose Routing-based On-Policy Distillation (ROPD), a novel realignment framework that models the divergence between aligned and compromised output probability distributions rather than fitting specific prompt templates. We conduct extensive experiments comparing ROPD against four state-of-the-art baselines across three datasets and three base models with varying alignment strengths. Our results demonstrate that when baseline defenses face template mismatches, often accompanied by severe degradation in downstream task performance. In contrast, ROPD substantially mitigates template-mismatch risks, maintaining superior robustness in both defense effectiveness and capability preservation. While our analysis indicates ROPD is not entirely immune to template shifts, its performance degradation is negligible compared to existing methods, establishing a new standard for robust LLM realignment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.