2607.26173v1 Jul 28, 2026 cs.LG

정렬 학습, 모델 유기체 및 시뮬레이션 모델 간의 공유된 지도 학습(SFT) 지식

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

Arthur Conmy
Arthur Conmy
Citations: 38
h-index: 4
Anton de la Fuente
Anton de la Fuente
Citations: 27
h-index: 2

정렬 학습, 모델 유기체 및 시뮬레이션 모델은 일반적으로 별개의 연구 분야로 취급됩니다. 그러나 이 세 가지 모두에서 진행되는 프로젝트들은 종종 동일한 근본적인 목표를 달성하기 위해 지도 미세 조정(SFT)을 활용합니다. 연구 프로젝트가 동일한 목표를 공유할 때, 한 영역에서 얻은 교훈이 다른 영역에 적용될 수 있는지 확인해야 합니다. 본 연구에서는 이러한 지식 전달의 세 가지 사례를 조사하며, 각 사례는 하나의 SFT 환경에서 개발된 교훈을 다른 환경에서 테스트합니다. 첫째, 정렬 학습에서 얻은 행동 일반화에 대한 교훈을 시뮬레이션 모델로 이전합니다. 'Teaching Claude Why'와 같이 특정 행동의 이유를 기반으로 훈련하면, 단순히 행동 예시만 사용하는 것보다 행동의 일반화 성능을 향상시킬 수 있습니다. 둘째, 모델 유기체에서 얻은 능력 보존에 대한 교훈을 Model-Spec Midtraining 정렬 환경으로 이전합니다. 학습 대상 모델이 아닌 다른 모델이 생성한 데이터(off-model 출력)를 사용하여 SFT를 수행하면, 훈련 과정에서 해당 모델의 능력이 손상될 수 있습니다. 그러나 훈련 데이터에 안전하고 적합한 on-model (및 on-policy) 데이터를 혼합하면 이러한 손상을 대부분 방지하면서도 목표 행동을 학습시킬 수 있습니다. 셋째, 모델 유기체에서 얻은 견고성(robustness)에 대한 교훈을 동일한 정렬 환경으로 이전합니다. 연구 결과, 후속적인 안전한 SFT는 정렬된 행동을 제거하는 동시에 능력을 보존할 수 있으며, 이는 능력 보존만으로는 이후 훈련에 대한 견고성을 보장하지 못한다는 것을 보여줍니다. 본 연구는 서로 다른 연구 분야 간의 SFT 지식을 이전함으로써 각 분야를 발전시킬 수 있음을 보여주며, 더 많은 연구자들이 자신의 전문 분야 외 영역에서 기술을 활용해야 할 것임을 시사합니다.

Original Abstract

Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a goal, we should test whether lessons learned from one area transfer to the other areas. We study three such transfers, each taking a lesson developed in one SFT setting and testing it in another. First, we port a lesson about behavior generalization from alignment training into toy models. Training on the reason for a behavior, as in Teaching Claude Why, can make the behavior generalize better than training on examples of the behavior alone. Second, we port a lesson about capability preservation from model organisms into the Model-Spec Midtraining alignment setting. SFT on outputs written by a model other than the student (off-model outputs) can damage capabilities when trained on. Mixing in benign on-model (and on-policy) data into our training can prevent most of this damage while still embedding the target behavior. Third, we port a lesson about robustness from model organisms into the same alignment setting. We find that follow-up benign SFT can erase the alignment behavior while preserving capabilities, showing that capability preservation alone does not ensure robustness to subsequent training. Our work illustrates how porting SFT lessons between different research fields can uplift them all, suggesting more researchers should borrow techniques from outside their own areas.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!