2606.16808v1 Jun 15, 2026 cs.AI

적응형 및 명시적인 안전: 대규모 추론 모델에서 잠재적 안전 인지 능력 활성화

Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models

Yuke Hu
Yuke Hu
Citations: 222
h-index: 8
Kehao Miao
Kehao Miao
Citations: 19
h-index: 3
Jiaxin Li
Jiaxin Li
Citations: 11
h-index: 1
Hongliang Chen
Hongliang Chen
Citations: 4
h-index: 1
Zhan Qin
Zhan Qin
Citations: 42
h-index: 3

대규모 추론 모델(LRM)은 복잡한 작업에서 뛰어난 성능을 보이지만, 정교한 탈옥 시도와 직접적인 유해 질문에 취약합니다. 이러한 취약점을 해결하기 위해 기존 연구는 안전 정렬을 위해 외부의 수동 데이터 주석에 크게 의존했습니다. 그러나 우리는 LRM이 원래 질문과 함께 자체 추론 과정을 다시 제시할 때, 고유하게 안전 위험을 식별할 수 있다는 것을 관찰했습니다. 이를 '잠재적 안전 인지(Latent Safety Awareness)'라고 명명합니다. 이러한 안전 인지 능력을 활용하기 위해, 먼저 지도 학습 미세 조정(SFT)을 사용하여 초기 추론 내용 이후에 안전 분석 및 지침을 활성화하는 데 필요한 안전 태그를 명시적으로 유도합니다. 동시에 일반적인 질문에는 표준 응답을 유지하여 적응형 트리거링을 보장합니다. 그 후, 직접 선호도 최적화(DPO)를 적용하여 안전 분석 및 지침의 정확성과 안정성을 더욱 향상시킵니다. 주목할 점은 두 단계의 학습에 필요한 모든 응답이 최적화되는 모델 자체에서 생성됩니다. (Safe Trigger) SFT와 DPO를 통해 실험 결과, 상당한 수준의 안전성 향상이 확인되었습니다. 예를 들어, DeepSeek-R1-Distill-Llama-8B 모델의 공격 성공률(ASR)은 유해 및 탈옥 벤치마크에서 각각 평균적으로 24.65%와 36.72% 감소했습니다. 마지막으로, Safe Trigger 방법은 일반적인 성능이나 사용자 경험에 거의 영향을 미치지 않습니다.

Original Abstract

While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, prior works depend heavily on external manual data annotation for safety alignment. However, we observe that LRMs can inherently identify safety risks when being re-presented with original queries alongside their own reasoning trajectories -- a capability we term Latent Safety Awareness. To leverage this safety awareness, we first employ Supervised Fine-Tuning (SFT) to explicitly induce safe tags to trigger safety analysis and guidance following the initial reasoning content for unsafe queries, while preserving standard responses for general queries to ensure adaptive triggering. Subsequently, we apply Direct Preference Optimization (DPO) to further enhance the correctness and stability of the safety analysis and guidance. Notably, responses required for both training stages are entirely generated by models being optimized. With (Safe Trigger) SFT and DPO, experimental results demonstrate significant safety enhancement. For example, the Attack Success Rate (ASR) of DeepSeek-R1-Distill-Llama-8B, on average, drops 24.65% and 36.72% on harmful and jailbreak benchmarks, respectively. Finally, our Safe Trigger method exerts almost no negative impact on general performance or user experience.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!