양성 샘플에 숨겨진 유해한 지침으로부터의 방어
Defending Against Harmful Supervision Hidden in Benign Samples
기존의 방어 기법은 유해한 콘텐츠가 명시적으로 추가된 후속 미세 조정 데이터에 효과적이지만, 악의적인 공격자는 유해한 지침을 양성 작업 내부에 숨길 수 있습니다. 본 연구에서는 유해한 질문-답변 쌍을 양성 학습 샘플 내에 포함시키는 '임베디드 공격(Embedded Attack)'을 제안하고, 대표적인 안전장치가 개별 예시 수준에서 이를 탐지하지 못하는 것을 보여줍니다. 이러한 문제를 해결하기 위해 우리는 토큰 수준의 정규화를 통해 DPO 스타일의 대비 목적 함수 설계를 SFT(Supervised Fine-Tuning)에 적용한 '듀얼 레퍼런스 SFT (DR-SFT)'를 제안합니다. DR-SFT는 데이터 필터링 이상의 수준에서 유해한 미세 조정을 완화할 수 있습니다.
Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead hide harmful supervision inside benign tasks. We propose Embedded Attack, where harmful QA pairs are embedded within benign training samples, and show that representative guardrails often fail to detect them at the example level. To address this, we propose Dual-Reference SFT (DR-SFT), which adapts DPO-style contrastive objective design to SFT through token-level regularization, mitigating harmful fine-tuning beyond coarse data filtering.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.