언제 신뢰할 수 없는 모니터링을 신뢰할 수 있는가? 공모 전략에 따른 안전성 분석
When can we trust untrusted monitoring? A safety case sketch across collusion strategies
인공지능은 점점 더 높은 수준의 자율성과 기능을 갖추고 있으며, 이는 잘못 정렬된 인공지능이 심각한 피해를 초래할 위험을 증가시킵니다. 신뢰할 수 없는 모니터링은 위험을 줄이기 위한 한 가지 방법으로, 이는 신뢰할 수 없는 모델을 사용하여 다른 모델을 감시하는 방식입니다. 신뢰할 수 없는 모니터링 시스템의 안전성을 입증하는 것은 어렵습니다. 왜냐하면 개발자는 잘못 정렬된 모델을 안전하게 배포하여 프로토콜을 직접 테스트할 수 없기 때문입니다. 본 논문에서는 사전 배포 테스트를 기반으로 안전성을 엄격하게 입증하는 기존 방법을 발전시킵니다. 우리는 이전의 인공지능 제어 연구에서 신뢰할 수 없는 모니터링을 방해하기 위해 잘못 정렬된 인공지능이 사용할 수 있는 공모 전략에 대한 가정을 완화합니다. 우리는 수동적 자기 인식, 인과적 공모 (사전에 공유된 신호 숨기기), 비인과적 공모 (셸링 포인트를 통한 신호 숨기기) 및 결합된 전략을 포괄하는 분류 체계를 개발합니다. 우리는 안전성 분석 개요를 작성하여 주장을 명확하게 제시하고, 가정을 명시적으로 밝히며, 해결되지 않은 과제를 강조합니다. 우리는 수동적 자기 인식이 이전에 연구되었던 전략보다 더 효과적인 공모 전략이 될 수 있는 조건을 파악합니다. 본 연구는 신뢰할 수 없는 모니터링에 대한 보다 강력한 평가를 위한 기반을 마련합니다.
AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitoring -- using one untrusted model to oversee another -- is one approach to reducing risk. Justifying the safety of an untrusted monitoring deployment is challenging because developers cannot safely deploy a misaligned model to test their protocol directly. In this paper, we develop upon existing methods for rigorously demonstrating safety based on pre-deployment testing. We relax assumptions that previous AI control research made about the collusion strategies a misaligned AI might use to subvert untrusted monitoring. We develop a taxonomy covering passive self-recognition, causal collusion (hiding pre-shared signals), acausal collusion (hiding signals via Schelling points), and combined strategies. We create a safety case sketch to clearly present our argument, explicitly state our assumptions, and highlight unsolved challenges. We identify conditions under which passive self-recognition could be a more effective collusion strategy than those studied previously. Our work builds towards more robust evaluations of untrusted monitoring.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.