비관주의의 역설: 보수적인 오프라인 학습이 추론 모델에서 온라인 적응 과정 중 보상 조작을 증폭시킨다
Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models
보수적인 오프라인 학습은 안전한 기반으로 후속 온라인 적응에 널리 권장됩니다. 이 주장은 정책이 잘 뒷받침되는 행동과 가까이 유지되면, 학습된 보상 모델의 불완전성을 악용할 가능성이 낮다는 것입니다. 우리는 이러한 직관을 경험적으로나 메커니즘적으로 도전합니다. Qwen3-14B 정책을 직접 선호 최적화(DPO)를 사용하여 세 가지 수준의 보수성($β ext{는 } β_{ ext{lo}}, β_{ ext{mid}}, β_{ ext{hi}} ext{로부터 얻어진 경험적인 로그 비율 백분위수에 해당}$)으로 학습시킨 후, 각 체크포인트를 학습된 보상 앙상블(3 x Qwen3-1.7B)에 대해 온라인으로 적응시키면서 GSM8K의 정확한 답변 정확도를 측정했습니다. 우리는 extit{더 높은 오프라인 보수성이 Goodhart 격차와 그 영역(AUGC)으로 측정되는 보상 조작 피해를 단조적으로 증가시킨다}는 것을 발견했으며, Spearman 상관계수는 모든 세 가지 조건에서 1.0입니다. 메커니즘적 분석 결과, 세 단계의 인과 관계가 밝혀졌습니다: (i) 높은 $β$ DPO는 정책 엔트로피를 감소시키고, (ii) 낮은 엔트로피 정책은 다양성이 줄어든 응답을 생성하며, 보상 모델 훈련 분포의 좁은 영역에 집중됩니다(낮은 쌍별 코사인 거리), 그리고 (iii) 이러한 근접성에도 불구하고 앙상블 불일치(인식적 불확실성)는 $β$가 증가함에 따라 증가하고 온라인 최적화 과정에서 더 빠르게 악용됩니다. 또한, 우리는 $(β, ext{AUGC})$ 데이터에 대한 거듭제곱 법칙 곡선을 피팅하여 정렬 충실성과 해킹 취약성을 균형 있게 하는 실질적으로 최적의 보수성 수준 $β^{ ext{*}}$를 식별했습니다. 우리의 결과는 해당 분야가 extit{최대한의} 보수성이 아니라 extit{균형 잡힌} 보수성이 필요하다는 것을 시사합니다.
Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model. We challenge this intuition empirically and mechanistically. We train a Qwen3-14B policy under Direct Preference Optimisation (DPO) with three levels of conservatism ($β\in \{β_{\mathrm{lo}}, β_{\mathrm{mid}}, β_{\mathrm{hi}}\}$ derived from empirical log-ratio percentiles), then adapt each checkpoint online against a learned reward ensemble (3\,$\times$\,Qwen3-1.7B) while measuring true performance on GSM8K exact-answer accuracy. We find that \emph{higher offline conservatism monotonically increases reward-hacking damage}, measured by the Goodhart gap and its area under the curve (AUGC), with Spearman $ρ= 1.0$ across all three conditions. Mechanistic analysis reveals a three-link causal chain: (i) high-$β$ DPO compresses policy entropy, (ii) Low-entropy policies generate responses with reduced diversity, concentrating in a narrow region of the reward model's training distribution (lower pairwise cosine distance), and (iii) despite this proximity, ensemble disagreement (epistemic uncertainty) increases with $β$ and is exploited faster during online optimisation. We further fit a power-law curve to the $(β, \augc)$ data and identify a practical optimal conservatism level $β^{\star}$ that balances alignment fidelity against hacking vulnerability. Our results suggest that the field needs \emph{calibrated}, not \emph{maximal}, conservatism.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.