2606.19168v1 Jun 17, 2026 cs.AI

안전한 데이터 그 이상: 정기적인 안전성 검토를 통한 사전 학습 단계 정렬

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection

Kaifeng Lyu
Kaifeng Lyu
Citations: 95
h-index: 5
Jinhan Li
Jinhan Li
Citations: 0
h-index: 0
K. Tang
K. Tang
Citations: 139
h-index: 3
Yihan Xu
Yihan Xu
Citations: 0
h-index: 0
Zhuorui Ye
Zhuorui Ye
Citations: 14
h-index: 2

대규모 언어 모델(LLM)의 더 깊은 수준의 안전성 정렬을 달성하기 위해, 최근 연구에서는 안전 개입을 사전 학습 단계에서 더 일찍 적용하는 방법을 모색하고 있습니다. 이는 주로 위험한 데이터를 필터링하거나 더 안전한 형태로 재작성하는 방식으로 이루어집니다. 우리는 사전 학습 단계에서의 정렬이 단순히 데이터를 안전하게 만드는 것 이상이어야 한다고 주장합니다. LLM은 겉보기에 무해한 지식과 능력을 결합하여 위험한 행동을 유발할 수 있습니다. 이에, 우리는 '안전 검토 사전 학습(Safety Reflection Pretraining)'이라는 사전 학습 단계 정렬 방법을 제안합니다. 이 방법은 언어 모델링에 자기 모니터링 기능을 직접 통합하기 위해 짧은 안전성 검토 내용을 정기적으로 사전 학습 데이터에 삽입하며, 이는 이후의 추가 훈련을 통해 강화됩니다. FineWeb-Edu 데이터를 사용하여 사전 학습된 17억 개의 파라미터를 가진 모델로 실험한 결과, 안전 검토 사전 학습은 안전 분류 정확도를 향상시키고, 추론 단계 및 미세 조정 공격의 성공률을 크게 감소시켰습니다. 실제 환경에서의 실험과 더불어, 우리는 명확한 안전 정의와 모델이 안전 데이터로부터 위험한 행동을 쉽게 일반화할 수 있는 추론 구조를 갖춘 완전하게 제어된 합성 환경인 MedSafetyWorld를 소개합니다. MedSafetyWorld에서 수행된 추가 분석은 안전 검토 사전 학습이 데이터 필터링 및 재작성 방식에 비해, 모델이 안전 데이터로부터 일반화된 위험한 행동을 하지 않도록 방지하는 데 더 효과적임을 명확하게 보여줍니다. 종합적으로 볼 때, 우리의 연구 결과는 사전 학습 단계의 정렬이 단순히 훈련 데이터를 안전하게 만드는 것뿐만 아니라, 모델이 안전한 데이터로부터 습득할 가능성이 있는 행동을 형성해야 함을 시사합니다.

Original Abstract

To achieve deeper safety alignment for large language models (LLMs), recent efforts have studied how to push safety interventions earlier into the pretraining stage, primarily by filtering unsafe data or rewriting it into safer forms. We argue that pretraining-stage alignment should go beyond making the data safe: LLMs may compose seemingly benign knowledge and capabilities into unsafe behaviors. To this end, we propose Safety Reflection Pretraining, a pretraining-stage alignment method which regularly inserts short safety reflections into pretraining corpora to integrate self-monitoring directly into language modeling, establishing a foundational capability that is subsequently reinforced by compatible post-training. Our experiments with 1.7B models pretrained on FineWeb-Edu show that Safety Reflection Pretraining improves safety classification accuracy and substantially reduces the success rates of inference-stage and finetuning attacks. Complementary to our real-world experiments, we also introduce a fully controlled synthetic environment, MedSafetyWorld, with a clear definition of safety and a reasoning structure under which models can easily generalize unsafe behaviors from safe data. Ablations in MedSafetyWorld further demonstrate a clear advantage of Safety Reflection Pretraining in preventing models from acting on unsafe behaviors generalized from safe data, compared with data filtering and rewriting. Taken together, our findings suggest that pretraining alignment should not only make the training data safe, but also shape the behaviors that models are likely to acquire from safe data.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!