내성적 결합: 자기 설명 훈련은 고정된 감독 하에서도 행동 변화를 유도한다
Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision
언어 모델(LM)을 학습시켜 예측에 대한 설명을 생성하도록 할 때, 표면적인 모방이 아닌 진정한 자기 성찰을 얻을 수 있을까? 우리는 LM이 입력 데이터의 어떤 특징이 자신의 행동에 영향을 미쳤는지 설명하도록 훈련될 때, 모델의 역설계된 행동을 사용하여 감독을 수행했습니다. 놀랍게도, 이전 체크포인트 버전 또는 다른 계열의 유사한 모델에서 파생된 고정된 역설계 설명을 기반으로 학습된 LM은 종종 자신의 현재 행동에 더 충실한 설명을 생성하는 경향이 있으며, 이는 훈련 대상의 행동보다 더 그렇습니다. 이러한 "내성적" 결합은 훈련 과정 동안 설명이 현재 행동과 충분히 상관관계를 유지할 때 발생하며, 이때 행동 자체가 변화합니다. 또한, 내성적 결합은 행동 변화를 추적한다는 것을 보여줍니다. 즉, 설명 학습이 다른 사후 훈련 목표와 동시에 제공될 때, 설명은 새로운 감독 없이도 이러한 변화를 추적합니다. 이 현상은 아첨 및 거부와 같은 다양한 작업에서 나타나며, 라벨 노이즈에 강건합니다. 전반적으로, 우리의 결과는 역설계 설명의 고정된 데이터 세트가 자기 성찰을 위한 확장 가능하고 일반화 가능한 사후 훈련 신호를 제공할 수 있음을 보여줍니다.
When does training language models (LMs) to generate explanations of their predictions yield faithful introspection, rather than superficial imitation? We study LMs trained to explain which features of their inputs influenced their behavior, using models' counterfactual behavior on modified inputs as supervision. Surprisingly, we find that LMs trained on fixed counterfactual explanations derived from earlier checkpoints of themselves, or even from behaviorally similar models in different families, frequently produce explanations more faithful to their own current behaviors than to those of their training targets. This "introspective" coupling between LM explanations and behaviors occurs when training explanations remain sufficiently correlated with current behaviors over the course of training, even as behaviors themselves shift. We also show that introspective coupling tracks behavior shifts: when explanation training is provided concurrently with other post-training objectives, explanations track those shifts without requiring updated supervision. This phenomenon appears in multiple tasks, including sycophancy and refusal, and is robust to label noise. Overall, our results show that even fixed datasets of counterfactual explanations can provide scalable and generalizable post-training signal for introspection.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.