확장 가능한 감독을 위한 보수성 조정
Calibrating Conservatism for Scalable Oversight
자율적인 계획 및 광범위한 환경 상호 작용이 가능한 에이전트 기반 인공지능 시스템은 근본적인 제어 문제를 야기합니다. 즉, 인간은 자신의 능력 범위를 초월할 수 있는 시스템에 대해 어떻게 의미 있는 감독을 유지할 수 있을까요? 확장 가능한 감독을 위한 기존 접근 방식은 복잡한 가정을 전제로 하거나, 주로 휴리스틱적이며, 통계적 보장을 제공하는 실용적인 방법이 부족합니다. 본 연구에서는 다양한 보조 점수 함수를 결합하여 보수적인 기준선에서 벗어나는 정도를 나타내는 페널티를 측정하는 '교정된 집단 감독(Calibrated Collective Oversight, CCO)'을 제안합니다. Attainable Utility Preservation 개념에서 영감을 받은 CCO는 집단적 보수성을 가능하게 합니다. 즉, 행동은 감독자의 우려 수준에 비례하여 페널티를 받으므로, 높은 효용을 제공하는 행동은 감독자가 이를 수용할 경우에도 선택되지만, 우려가 누적될 때만 무시됩니다. CCO는 Conformal Decision Theory를 사용하여 이 보수성을 온라인으로 조정하며, 원치 않는 결과가 사용자가 지정한 목표 임계값 아래에 유지되도록 유한 시간 내에 보장하고, 분포에 대한 가정을 하지 않습니다. 수정된 SWE-bench 데이터셋에서, 약한 감독자는 적대적으로 잘못 정렬된 강력한 에이전트를 성공적으로 제약했으며, MACHIAVELLI 환경에서는 CCO가 윤리적 위반을 크게 줄이는 동시에 보상을 유지했습니다. 두 가지 설정 모두에서, 실제 위반율은 이론적으로 예측된 목표 값과 밀접하게 일치했습니다.
Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain meaningful oversight of systems that may exceed their own capabilities? Existing approaches to scalable oversight rely on complex assumptions, remain largely heuristic, or lack practical methods for sequential settings with statistical guarantees. We introduce Calibrated Collective Oversight (CCO), which aggregates diverse auxiliary scoring functions into a penalty measuring deviation from a conservative baseline. Inspired by Attainable Utility Preservation, CCO enables collective conservatism: actions face a penalty proportional to overseer concern, so high-utility actions are still selected when overseers find them unobjectionable and overridden only when concern accumulates. CCO calibrates this conservatism online using Conformal Decision Theory, ensuring that undesirable outcomes remain below a user-specified target threshold with finite-time bounds and no distributional assumptions. On a modified version of SWE-bench, weaker overseers successfully constrain an adversarially misaligned stronger agent; on MACHIAVELLI, CCO substantially reduces ethical violations while preserving reward. In both settings, empirical violation rates closely match the specified targets, as predicted by the theory.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.