2608.09542v1 Aug 10, 2026 cs.LG

이중 적대적 안전 정렬: LRM (Large Reasoning Models)에서 내재적인 위협 이해 능력 함양

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Shaopeng Fu
Shaopeng Fu
Citations: 31
h-index: 2
Qinbo Zhang
Qinbo Zhang
Citations: 11
h-index: 2

LLM(Large Language Models, 대규모 언어 모델)은 복잡한 작업에서 놀라운 성과를 보이지만, 여전히 유해한 프롬프트에 취약하여 안전하지 않은 결과를 초래할 수 있습니다. 최근에는 직접적인 거부 또는 안전 관련 설명을 사용하여 LLM을 정렬하는 방법들이 제시되었지만, 이러한 방법들은 종종 특정 프롬프트 패턴에 집중하며, 근본적인 공격 메커니즘까지 고려하지 않습니다. 결과적으로, 이러한 패턴 기반 정렬은 다양한 탈옥 시도(jailbreak)에 대한 일반화 능력이 부족하여 적대적 강건성(adversarial robustness)과 추론 능력(reasoning utility)을 저해합니다. 본 논문에서는 LLM이 명시적으로 적대적인 메커니즘을 분해함으로써 내재적인 안전 지식을 습득하도록 하는 이중 적대적 프레임워크인 AdvSafe를 제안합니다. 이는 패턴에 의존적인 정보를 넘어, 추론 능력을 손상시키지 않으면서 강력한 인지적 방어를 가능하게 합니다. 우리 시스템은 두 단계의 적대적 게임으로 작동합니다. 첫째, 적대적 생성 단계에서는 자율 에이전트가 기만적인 탈옥 프롬프트를 동적으로 생성하며, 강력한 모델을 공격하기 위한 전략을 조정합니다. 둘째, 적대적 추출 단계에서는 공격에 성공한 모델이 인지적 반격을 수행합니다. 각 성공적인 탈옥 시도에 대해, 모델은 공격의 원리를 밝히고, 그러한 프롬프트를 어떻게 식별하고 완화할 수 있는지 설명합니다. 이러한 이중 적대적 프로세스를 통해 풍부하고 일반화 가능한 안전 지식을 담은 간결한 추론 데이터셋을 구축합니다. 이 데이터셋으로 학습된 모델들은 내재적인 위협 이해를 통해 안전 정렬 능력을 획득하게 됩니다. 실험 결과, AdvSafe로 정렬된 LLM은 단 1,000개의 생성된 샘플만 사용했을 때에도 기존의 방법들보다 훨씬 강력한 탈옥 저항성을 보이며, 추론 능력의 거의 눈에 띄지 않는 손실만을 경험했습니다. 또한, AdvSafe는 일반화되지 않은 프롬프트에 대한 강건성도 향상시켜, 안전 지식 학습을 통해 우수한 강건성-유틸리티 균형을 달성하고 기존 공격 패턴 너머까지 일반화될 수 있음을 보여줍니다.

Original Abstract

Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direct refusals or safety rationales, yet often focus on prompt patterns rather than intrinsic attack mechanisms. As a result, these pattern-centric alignments struggle to generalize across diverse jailbreaks, compromising adversarial robustness and reasoning utility. We propose AdvSafe, a dual-adversarial framework that enables LRMs to internalize unsafety knowledge by explicitly deconstructing adversarial mechanisms. This moves beyond pattern-dependent traces, fostering robust cognitive defense without compromising reasoning utility. Our pipeline operates via a two-phase adversarial game. First, in adversarial synthesis, an autonomous agent dynamically crafts deceptive jailbreak prompts, adapting its strategies to breach a strong teacher model. Second, in adversarial extraction, the breached teacher executes a cognitive counter-attack. For every successful jailbreak, the teacher unmasks the camouflage, explaining why the attack succeeds and how such prompts can be identified and mitigated. This dual-adversarial process yields a compact reasoning dataset capturing rich, generalizable unsafety knowledge. Student models trained on this dataset implicitly acquire safety alignment through intrinsic threat comprehension. Experiments show that with only 1K synthesized samples, AdvSafe-aligned LRMs achieve significantly stronger jailbreak robustness than existing baselines, with almost no utility degradation. Furthermore, AdvSafe improves robustness against out-of-distribution prompts, demonstrating that learning unsafety knowledge enables a superior robustness-utility trade-off and generalizes beyond seen attack patterns.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!