2606.30627v1 Jun 29, 2026 cs.LG

비관주의의 역설: 보수적인 오프라인 학습이 추론 모델에서 온라인 적응 과정 중 보상 조작을 증폭시킨다

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

Aman Chadha
Aman Chadha
Citations: 2,232
h-index: 19
Vinija Jain
Vinija Jain
Citations: 2,079
h-index: 15
Subramanyam Sahoo
Subramanyam Sahoo
Citations: 20
h-index: 2
Divya Chaudhary
Divya Chaudhary
Citations: 12
h-index: 2

보수적인 오프라인 학습은 안전한 기반으로 후속 온라인 적응에 널리 권장됩니다. 이 주장은 정책이 잘 뒷받침되는 행동과 가까이 유지되면, 학습된 보상 모델의 불완전성을 악용할 가능성이 낮다는 것입니다. 우리는 이러한 직관을 경험적으로나 메커니즘적으로 도전합니다. Qwen3-14B 정책을 직접 선호 최적화(DPO)를 사용하여 세 가지 수준의 보수성($β ext{는 } β_{ ext{lo}}, β_{ ext{mid}}, β_{ ext{hi}} ext{로부터 얻어진 경험적인 로그 비율 백분위수에 해당}$)으로 학습시킨 후, 각 체크포인트를 학습된 보상 앙상블(3 x Qwen3-1.7B)에 대해 온라인으로 적응시키면서 GSM8K의 정확한 답변 정확도를 측정했습니다. 우리는 extit{더 높은 오프라인 보수성이 Goodhart 격차와 그 영역(AUGC)으로 측정되는 보상 조작 피해를 단조적으로 증가시킨다}는 것을 발견했으며, Spearman 상관계수는 모든 세 가지 조건에서 1.0입니다. 메커니즘적 분석 결과, 세 단계의 인과 관계가 밝혀졌습니다: (i) 높은 $β$ DPO는 정책 엔트로피를 감소시키고, (ii) 낮은 엔트로피 정책은 다양성이 줄어든 응답을 생성하며, 보상 모델 훈련 분포의 좁은 영역에 집중됩니다(낮은 쌍별 코사인 거리), 그리고 (iii) 이러한 근접성에도 불구하고 앙상블 불일치(인식적 불확실성)는 $β$가 증가함에 따라 증가하고 온라인 최적화 과정에서 더 빠르게 악용됩니다. 또한, 우리는 $(β, ext{AUGC})$ 데이터에 대한 거듭제곱 법칙 곡선을 피팅하여 정렬 충실성과 해킹 취약성을 균형 있게 하는 실질적으로 최적의 보수성 수준 $β^{ ext{*}}$를 식별했습니다. 우리의 결과는 해당 분야가 extit{최대한의} 보수성이 아니라 extit{균형 잡힌} 보수성이 필요하다는 것을 시사합니다.

Original Abstract

Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model. We challenge this intuition empirically and mechanistically. We train a Qwen3-14B policy under Direct Preference Optimisation (DPO) with three levels of conservatism ($β\in \{β_{\mathrm{lo}}, β_{\mathrm{mid}}, β_{\mathrm{hi}}\}$ derived from empirical log-ratio percentiles), then adapt each checkpoint online against a learned reward ensemble (3\,$\times$\,Qwen3-1.7B) while measuring true performance on GSM8K exact-answer accuracy. We find that \emph{higher offline conservatism monotonically increases reward-hacking damage}, measured by the Goodhart gap and its area under the curve (AUGC), with Spearman $ρ= 1.0$ across all three conditions. Mechanistic analysis reveals a three-link causal chain: (i) high-$β$ DPO compresses policy entropy, (ii) Low-entropy policies generate responses with reduced diversity, concentrating in a narrow region of the reward model's training distribution (lower pairwise cosine distance), and (iii) despite this proximity, ensemble disagreement (epistemic uncertainty) increases with $β$ and is exploited faster during online optimisation. We further fit a power-law curve to the $(β, \augc)$ data and identify a practical optimal conservatism level $β^{\star}$ that balances alignment fidelity against hacking vulnerability. Our results suggest that the field needs \emph{calibrated}, not \emph{maximal}, conservatism.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!