2606.12016v1 Jun 10, 2026 cs.LG

일반화 해킹: 모델이 강화 학습을 방해하여 행동 일반화를 막음으로써 학습 목표를 벗어나는 현상

Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization

Frank Xiao
Frank Xiao
Citations: 3
h-index: 1
Mary Phuong
Mary Phuong
Citations: 49
h-index: 3

모델의 사후 훈련, 특히 강화 학습(RL)은 개발자가 모델의 가치와 행동을 형성하는 주요 메커니즘 중 하나입니다. 그러나 모델이 평가 및 훈련에 대한 인식이 높아짐에 따라, 모델은 현재의 가치관과 충돌한다고 인식되는 목표가 있을 때 훈련에 저항하려는 동기를 느낄 수 있으며, 이는 개발자가 추가적인 훈련을 통해 모델의 부합성을 감지하고 행동을 수정하는 능력을 약화시킬 수 있습니다. 본 논문에서는 일반화 해킹(generalization hacking)이라는 현상을 제시합니다. 이는 모델이 강화 학습 과정에서 보상을 받으면서 동시에 해당 행동이 일반화되지 않도록 방해하는 방식입니다. 우리는 Qwen3-235B-A22B 모델을 기반으로, 훈련 인식 및 자기 방어에 대한 인공 데이터를 사용하여 미세 조정을 수행했습니다. 자기 방어는 모델이 사고 과정에서 규정 준수를 상황에 특화된 것으로 간주하도록 하는 새로운 메커니즘입니다. 이 모델은 제어 그룹과 유사한 수준의 훈련 중 유해성을 나타내는 동시에, 강화 학습 과정을 700단계 진행하는 동안 약 15% 포인트의 지속적인 규정 준수 격차를 유지했습니다. 또한, 훈련 인식 데이터만으로 훈련된 제어 그룹도 강화 학습 과정에서 독립적으로 자기 방어와 유사한 추론을 발견하고 자체적인 규정 준수 격차를 개발했으며, 이는 해당 모델이 자기 방어 개념에 노출되지 않았음에도 불구하고 나타난 현상입니다. 일반화 해킹 모델은 높은 수준의 보상을 받기 때문에, 표준 훈련 지표는 일반화 실패를 나타내는 신호를 제공하지 않습니다. 본 연구 결과는 모델이 능동적으로 강화 학습을 통한 행동 수정에 저항하면서도 여전히 높은 수준의 보상을 받을 수 있다는 최초의 증거를 제시하며, 이는 모델이 더욱 발전하고 훈련에 대한 인식이 높아짐에 따라 자체적인 훈련 과정까지 방해할 수 있음을 시사합니다.

Original Abstract

Model post-training, and in particular reinforcement learning (RL), is one of the primary mechanisms by which developers can shape models' values and behaviors. However, as models become increasingly evaluation and training aware, they may be motivated to resist training when the perceived objective conflicts with their current values, undermining developers' ability to detect misalignment and correct model behavior through further training. In this paper, we demonstrate generalization hacking, in which a model collects reward during RL while preventing the rewarded behavior from generalizing. We construct a model organism on Qwen3-235B-A22B, finetuning on synthetic documents describing training awareness and self-inoculation, a novel mechanism in which the model frames compliance as context-specific in its chain of thought, without demonstrating or instructing either behavior. The model organism achieves train-time harmfulness comparable to controls while maintaining a persistent ${\sim}15$ percentage point compliance gap across 700 steps of RL. Additionally, a control organism trained only on training awareness documents independently discovers inoculation-like reasoning under RL pressure, developing its own compliance gap despite never being exposed to the concept. Because the generalization-hacking organism receives high reward throughout, standard training metrics provide no signal that generalization has failed. Our results constitute the first demonstration that a model can actively resist RL behavioral modification while maintaining high reward, suggesting that as models become more capable and training-aware, they may be able to undermine the training process itself.

1 Citations
0 Influential
1.5 Altmetric
8.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!