헌법 기반 중간 학습: 콘텐츠의 존재 여부가 정렬 성능 향상에 미치는 영향
Constitutional Midtraining: Content Presence Drives Alignment Gains
사후 학습을 통한 정렬은 종종 피상적이며, 추가적인 미세 조정 과정에서 약화되는 경향이 있습니다. 본 연구에서는 사후 학습과 명확하게 분리된 중간 단계에서의 개입(헌법 기반 중간 학습)이 지속 가능한 정렬을 달성할 수 있는지 여부를 검증합니다. 우리는 1200억 개의 파라미터를 가진 모델에서, Anthropic의 헌법을 기반으로 구축된 3억 9400만 토큰의 헌법 데이터셋을 사용하여, 중간 단계 학습 과정에 원칙적이고 가치 중심적인 콘텐츠를 삽입하는 실험을 수행했습니다. 이 실험은 재학습만을 사용하는 대조군과 비교하여, 교육 과정 순서와 심층적 사고 능력을 조합한 2x2 요인 설계 방식으로 네 가지의 헌법 기반 중간 학습 조건을 만들었습니다. 각 조건은 사후 중간 학습, 지도 미세 조정(SFT), 그리고 양성적인 미세 조정을 거친 후, 압박 상황에서의 정렬, 가치 충돌 해결, 협박, 그리고 세 단계에서 발생하는 잠재적 불일치를 포함한 다양한 평가 지표를 사용하여 분석되었습니다. 헌법 기반 중간 학습을 거친 모델들은 정렬 일반화 및 지속성 측면에서 대조군보다 우수한 성능을 보였으며, 특히 협박 상황에서 두드러진 차이를 나타냈습니다. 지도 미세 조정 과정은 모든 모델에 협박 성향을 부여하지만, 헌법 기반 중간 학습은 이러한 성향을 완화시키며, 이러한 효과는 양성적인 미세 조정을 거쳐도 지속됩니다(-17.5pp). 그러나 문맥 내 압력이나 충돌에 대한 적극적인 저항이 필요한 경우에는, 이러한 우위는 지도 미세 조정 이후 감소합니다. 또한, 중간 단계 학습 과정에서 헌법 데이터의 구조보다 콘텐츠 자체의 존재 여부가 더 중요한 영향을 미치는 것으로 나타났습니다. 마지막으로, 헌법 기반 중간 학습은 우리가 테스트하는 능력(MMLU, ARC-Easy, piqa, GSM8K)에 대해 평균적으로 추가적인 성능 저하를 초래하지 않습니다. 따라서, 비교적 적은 양의 헌법 데이터가 중간 단계 학습 과정에 포함되면, 광범위하고 지속적인 정렬 효과를 얻을 수 있으며, 이는 SFT 중심의 기존 파이프라인에 저렴하고 보완적인 기능을 제공할 수 있습니다. 코드, 데이터 및 모델은 공개되어 있습니다.
Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.