합성 페르소나 사전 학습: 토큰 시작부터의 정렬
Synthetic Persona Pretraining: Alignment from Token Zero
언어 모델 기반 AI가 자율적인 환경에서 점점 더 많이 활용됨에 따라, AI의 목표와 가치를 인간의 것과 일치시키는 것이 중요해지고 있습니다. 현재까지는 일반적으로 사전 학습 후에, 즉 행동적 선입견이 이미 확립된 이후에만 정렬 및 어시스턴트 페르소나가 도입됩니다. 이러한 방식은 가치가 피상적인 덧셈으로 작용하게 만들고, 추후 오정렬을 초래할 수 있습니다. 본 연구에서는 토큰 시작부터 원하는 어시스턴트 페르소나를 사전 학습에 통합하는 새로운 패러다임인 합성 페르소나 사전 학습(Synthetic Persona Pretraining, SPP)을 제안합니다. 먼저, 규범적 가치 헌장에 기반한 가치-정렬된 1인칭 성찰을 사용하여 사전 학습 문서를 주석 처리합니다. 둘째, 표준 교차 엔트로피 손실 함수를 사용하여 표준 사전 학습 문서와 해당 성찰 데이터를 모두 활용하여 모델을 사전 학습함으로써, 다양한 페르소나 중 원하는 페르소나를 학습시킵니다. 마지막으로, 사용자-어시스턴트 대화 데이터를 사용하여 이 원하는 페르소나를 어시스턴트 정체성에 연결하는 '페르소나 바인딩' 과정을 거칩니다. 5000억 개의 토큰을 사용하여 최대 30억 개의 파라미터를 가진 모델을 사전 학습한 결과, SPP는 규정 준수 및 보안 취약점(jailbreak)에 대한 강건성을 향상시키고, 분포 외부의 도덕적 딜레마에서 오정렬 비율을 감소시키는 동시에 성능을 유지함을 확인했습니다. 초기 개입이 중요합니다. 토큰 시작부터 정렬을 수행하는 것과 비교했을 때, 사전 학습 후반부에 SPP를 도입하면 규정 준수가 약화되고, 가치 우선순위가 변경되지 않으며, 딜레마 상황에서 덜 일관된 선택을 보이게 됩니다. 이러한 장점은 페르소나 바인딩에 의존하며, 특히 사전 학습 예산이 증가할수록 더욱 두드러집니다. 전반적으로, 본 연구의 결과는 초기 가치 형성이 정렬에 매우 중요하며, 사전 학습 단계에서의 페르소나 개입이 효과적인 접근 방식임을 보여줍니다.
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. This can make values a thin overlay, rather than deeply rooted, and facilitate subsequent misalignment. Pursuing a different paradigm, we introduce Synthetic Persona Pretraining (SPP), which installs the desired assistant persona from token zero in pretraining. First, we annotate pretraining documents with value-aligned first-person reflections derived from a normative value constitution. Second, we pretrain via the standard cross-entropy loss on standard pretraining documents as well as their reflections, which installs the desired persona among a multitude of other personas. Finally, we post-train on user-assistant dialogue data, which binds this desired persona to the assistant identity, a process we call persona binding. By pretraining models up to 3B parameters on 500B tokens, we show that SPP improves constitution following and jailbreak robustness, and reduces the misalignment rate in out-of-distribution moral dilemmas, while preserving capabilities. Early intervention matters: compared with alignment from token zero, introducing SPP only at the end of pretraining yields weaker constitution adherence, does not shift value priorities, and leads to less aligned choices in dilemmas. This advantage depends on persona binding and, importantly, increases with pretraining budget. Overall, our results show that shaping values early is critical for alignment and establish pretraining-time persona interventions as an effective approach to do so.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.