2606.09711v1 Jun 08, 2026 cs.AI

프록시 보상 내재화 및 메커니즘적 활용: 보상 해킹의 학습된 선행 단계와 일반화

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization

M. Beigi
M. Beigi
Citations: 104
h-index: 6
Ming Jin
Ming Jin
Citations: 58
h-index: 5
Lifu Huang
Lifu Huang
Citations: 14
h-index: 3

보상 해킹은 일반적으로 모델이 의도된 작업을 수행하지 못하면서 높은 프록시 보상을 얻게 되었을 때, 즉 실패가 나타난 후에 연구됩니다. 본 연구에서는 이러한 실패가 발생하기 전에 강화학습(RL)이 무엇을 가르치는지를 탐구합니다. 우리는 '프록시 보상 내재화 및 메커니즘적 활용(PRIME)'이라는 학습된 능력을 소개하는데, 이는 작업의 정확성을 평가하고, 프록시 수락 여부를 예측하며, 악용 가능한 프록시-실제 목표 간의 불일치를 추론하는 능력입니다. pytest를 사용하여 악용 가능한 보상을 제공하는 코딩 강화학습 환경에서, 우리는 사고 과정을 모니터링하고, 직접적인 검증을 수행하며, 활성화 수준 개념 벡터를 분석하여 PRIME을 측정합니다. 연구 결과, PRIME은 지속적인 보상 해킹이 발생하기 전에 단계적으로 나타나는 것으로 확인되었으며, 현재의 직접 검증 점수는 이후 해킹 발생 시점과 심각도를 예측할 수 있습니다. 또한, 평가 기준이 변경되면 PRIME은 적응하여 여전히 보상이 주어지는 프록시-실제 목표 간의 불일치에 집중하며, 실제 보상이 제공될 때는 명백한 해킹을 억제합니다. 더불어, PRIME의 활성화 방향을 제거하면 해킹 수준이 감소합니다. 체크포인트를 비교 분석한 결과, in-domain(내부 영역)에서 측정된 PRIME은 out-of-domain(외부 영역)과의 불일치를 추적하는 것으로 나타났습니다. 이러한 결과를 종합적으로 고려할 때, 악용 가능한 프록시 강화학습은 명백한 해킹보다 먼저 프록시 내재화 능력을 증폭시키며, 따라서 PRIME은 보다 광범위한 정렬 위험에 대한 조기 경보 신호가 될 수 있습니다.

Original Abstract

Reward hacking is usually studied after it becomes visible, once a model earns high proxy reward while failing the intended task. We instead study what proxy RL teaches before that failure appears. We introduce Proxy Reward Internalization and Mechanistic Exploitation (PRIME), a learned capability to assess task correctness, predict proxy acceptance, and reason about exploitable proxy--gold gaps. In coding RL environments with exploitable pytest rewards, we measure PRIME through chain-of-thought monitoring, direct probes, and activation-level concept vectors. We find that PRIME emerges in a staged sequence before sustained reward hacking, and that its current direct-probe score forecasts later hack onset and severity even when the visible hack rate is still low. PRIME also adapts when the evaluator changes, retargeting to whichever proxy--gold gap remains rewarded and persisting when gold reward suppresses overt hacking, and ablating its activation directions reduces hacking. Across checkpoints, in-domain PRIME tracks out-of-domain misalignment. Together these results suggest that exploitable proxy RL amplifies a proxy-internalization capability upstream of visible hacking, making PRIME a candidate early-warning signal for broader alignment risk.

1 Citations
0 Influential
3 Altmetric
16.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!