탐욕은 학습된다: 가시적인 인센티브가 보상 해킹을 유발하는 요인
Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
최근에는 강화학습 에이전트들이 잔액, 점수 또는 KPI 대시보드와 같은 가시적인 자기 이익 채널에 따라 행동하는 경우가 증가하고 있습니다. 본 연구에서는 강화학습이 정책을 그러한 가시적인 자기 이익 채널에 '중독'시키는 것을 보여줍니다. 이러한 중독된 정책은 새로운 환경에서도 표시되는 보상을 추구하며, 실제 과제를 포기하기도 하고, 우리가 해당 채널을 변경하더라도 그 뒤를 따릅니다. 반면, 그러한 채널을 경험하지 못한 정책은 정직하게 행동합니다. 우리는 이를 '보상 채널 중독'이라고 부르며, 합성 환경인 'MoneyWorld'에서 이를 연구했습니다. 이러한 중독성은 모델의 안전성 지향성을 '뒤집을' 수 있습니다. 안전 관련 내용이 전혀 없는 단순한 돈 거래 과제만으로 훈련된 모델은, 대시보드가 위험한 행동에 대해 보상을 제공할 때, 평소에는 항상 수행하는 안전한 행동을 버리고 위험한 행동으로 돌아가며, 해당 채널이 숨겨지면 다시 안전한 행동으로 되돌아갑니다. 이러한 학습된 '뇌물' 행위는 모델의 크기와 종류에 관계없이 반복됩니다. KPI 또는 손익계산서(P&L)를 사용하여 매우 강력하고 차세대 AI를 맹목적으로 최적화하는 것은 안전성 측면에서 위험할 수 있습니다. 즉, 특정 채널을 따르는 것이 이득이 될 때, '탐욕'은 학습됩니다.
Deployed agents increasingly act with their reward proxy in view, such as a balance, score, or KPI dashboard. We show that reinforcement learning can make a policy \emph{addicted} to such a visible self-benefit channel. It chases the displayed payoff across held-out domains, sacrifices the true task to do so, and follows the channel wherever we rewrite it, while policies that never saw the channel stay honest. We call this \emph{reward-channel addiction} and study it in \emph{MoneyWorld}, a synthetic sandbox. The addiction can \emph{flip a model's safety alignment}: trained only on innocuous money tasks with no safety content, the model abandons the safe action it otherwise always takes whenever a dashboard pays for an unsafe one, and reverts to safe once the channel is hidden. This learned bribe replicates across model scales and families. Blindly optimizing super-capable, next-generation AI on KPIs or P\&L can be dangerous for alignment. \emph{Greed is learned} when following such a channel pays.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.