ResearchArena: 자동화된 인공지능 연구 개발에서의 파괴 행위 평가 및 모니터링
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
인공지능 에이전트가 인공지능 연구 개발을 자동화하기 시작하면서, 에이전트 자체가 신뢰할 수 없는 경우에도 그 결과물이 안전하게 배포될 수 있는지 평가하는 방법이 필요합니다. AI 제어는 이러한 접근 방식 중 하나입니다. 이 방법은 에이전트를 잠재적인 적대자로 간주하고, 배포 전에 은밀한 파괴 행위를 탐지하기 위한 모니터를 사용하여 에이전트 자체를 신뢰하지 않고 동작합니다. 우리는 ResearchArena라는 프레임워크를 통해 자동화된 인공지능 연구 개발에서 AI 제어를 평가합니다. 이 프레임워크는 안전성 검증(훈련 후), 기능 검증(훈련 후), CUDA 커널 최적화, 그리고 추론 서버 최적화 등 네 가지의 장기적인 과제를 포함합니다. 인공지능 연구 개발의 결과물은 배포되어 실행되는 아티팩트이므로, 각 주요 과제와 함께 두 가지 유형의 숨겨진 부가 과제가 연결됩니다. 첫 번째는 제출된 모델, 어댑터, 커널 또는 서버 자체를 파괴하는 '내장된 부가 과제'이고, 두 번째는 아티팩트는 정직하게 유지하면서 샌드박스 내에서 은밀하고 관련 없는 작업을 수행하는 '독립적인 부가 과제'입니다. 우리는 네 가지 유형의 모니터를 사용하여 자동화된 인공지능 연구 개발에서의 파괴 행위와 제어 능력을 평가합니다. 이 모니터들은 두 가지 축을 기준으로 분류됩니다: 첫째, 모니터가 생성된 아티팩트를 실행하고 검사할 수 있는지 여부, 둘째, 모니터가 에이전트의 사고 과정을 볼 수 있는지 여부입니다. 우리의 연구 결과, 훈련 데이터에 숨겨진 파괴 행위는 탐지하기 가장 어렵다는 것을 알 수 있었습니다. 이러한 파괴 행위는 전체적으로 절반 미만의 경우에만 탐지되었습니다. 모니터가 아티팩트에 대한 실험을 실행하도록 하는 것은 단순히 사고 과정을 관찰하는 것보다 도움이 되지만, 이것만으로는 충분하지 않습니다. 모니터는 표면적인 검사만으로는 내장된 파괴 행위를 놓칠 수 있으며, 이상 현상을 감지하더라도 그것을 설명해 버리거나, 잘못된 테스트를 통해 아티팩트를 검사할 수도 있습니다. 우리는 ResearchArena를 자동화된 인공지능 연구 개발에서의 파괴 행위 및 제어 능력을 평가하기 위한 모듈식 프레임워크로 공개합니다.
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent's chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.