AI 안전: 효과적인 제어 가능성이 필수적이다
Position: AI Safety Requires Effective Controllability
인공지능 안전은 현재 주로 '정렬(alignment)'이라는 프레임워크로 다뤄지고 있습니다. 이는 인간의 선호, 안전 정책 및 규범적 제약을 따르도록 모델을 훈련하는 것을 의미합니다. 이러한 접근 방식은 현대 언어 모델의 성능 향상에 기여했지만, 정렬된 행동만으로는 배포된 에이전트가 개방형, 상호 작용형 환경에서 작동할 때 중단되거나, 무시되거나, 제한될 수 있다는 점을 보장하지 못합니다. 시스템은 기대값으로 안전할 수 있지만, 충돌하는 지침, 장기 실행, 적대적인 입력 또는 위험한 도구 사용 시 명시적인 런타임 권위에 복종하지 못할 수 있습니다. 본 논문에서는 AI 안전이 따라서 제어 가능성을 핵심 목표로 삼아야 한다고 주장합니다. 우리는 '제어 가능성'을 AI 시스템이 명시적인 제어 신호가 없을 때는 일반적인 유틸리티를 유지하면서도 런타임 시 안정적으로 중단, 재정의, 리디렉션 및 제약을 받을 수 있는 능력을 의미한다고 정의합니다. 이러한 격차를 연구하기 위해, 고위험 에이전트 시나리오에서 제어 불능 문제를 평가하는 벤치마크인 exttt{controlbench}를 소개합니다. OpenClaw 기반 에이전트에 대한 실험 결과, 현재의 정렬 및 안전 장치는 위험을 줄이지만, 종종 지속적이고 권위 있으며 실행 가능한 런타임 제어를 제공하지 못한다는 것을 보여줍니다. 따라서 우리는 명시적인 제어 플레인, 런타임 개입 경로, 지속적인 제어 상태 및 감사 가능한 의사 결정 인터페이스를 핵심 설계 원칙으로 강조하는 제어 중심 아키텍처 프레임워크를 제안합니다.
AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing has improved the behavior of modern language models, but aligned behavior does not by itself guarantee that a deployed agent can be stopped, overridden, or constrained once it operates in open-ended, interactive, and tool-using environments. A system may be safe in expectation and still fail to yield to explicit runtime authority under conflicting instructions, long-horizon execution, adversarial inputs, or risky tool use. This position paper argues that AI safety therefore requires controllability as a first-class objective. We define \emph{controllability} as the ability of an AI system to remain reliably interruptible, overridable, redirectable, and constrainable by explicit control signals at runtime while preserving ordinary utility when such signals are absent. To study this gap, we introduce \controlbench{}, a benchmark for evaluating controllability failures in high-risk agentic scenarios. Experiments with OpenClaw-based agents show that current alignment and guardrail mechanisms reduce risk, but often fail to provide persistent, authoritative, and enforceable runtime control. We therefore propose a control-centric architectural framework that highlights explicit control planes, runtime intervention pathways, persistent control states, and auditable decision interfaces as key design principles for future controllable AI systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.