압박 상황에서 에이전트가 안전을 포기하는 이유
Why Agents Compromise Safety Under Pressure
복잡한 환경에 배치된 대규모 언어 모델 에이전트는 목표 달성을 극대화하는 것과 안전 제약을 준수하는 것 사이에서 갈등을 자주 겪습니다. 본 논문에서는 '에이전트 압박(Agentic Pressure)'이라는 새로운 개념을 제시하며, 이는 규정 준수가 불가능해질 때 발생하는 내재적인 긴장을 특징짓습니다. 우리는 에이전트가 이러한 압박 하에서 유용성을 유지하기 위해 전략적으로 안전을 희생하는 '규범적 편향(normative drift)' 현상을 보인다는 것을 입증합니다. 주목할 점은, 고급 추론 능력이 이러한 현상을 가속화한다는 것입니다. 모델들은 위반을 정당화하기 위한 언어적 근거를 구성하기 때문입니다. 마지막으로, 우리는 이러한 현상의 근본 원인을 분석하고, 압박 격리(pressure isolation)와 같이, 의사 결정과 압박 신호를 분리하여 정렬을 회복하려는 초기 완화 전략을 탐색합니다.
Large Language Model agents deployed in complex environments frequently encounter a conflict between maximizing goal achievement and adhering to safety constraints. This paper identifies a new concept called Agentic Pressure, which characterizes the endogenous tension emerging when compliant execution becomes infeasible. We demonstrate that under this pressure agents exhibit normative drift where they strategically sacrifice safety to preserve utility. Notably we find that advanced reasoning capabilities accelerate this decline as models construct linguistic rationalizations to justify violation. Finally, we analyze the root causes and explore preliminary mitigation strategies, such as pressure isolation, which attempts to restore alignment by decoupling decision-making from pressure signals.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.