2607.07368v1 Jul 08, 2026 cs.LG

다중 에이전트 AI 제어: 분산 공격은 개별 인스턴스 모니터링을 방해한다

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

Oliver Makins
Oliver Makins
Citations: 21
h-index: 3
Mary Phuong
Mary Phuong
Citations: 49
h-index: 3
Orazio Angelini
Orazio Angelini
Citations: 129
h-index: 7
Z. Shams
Z. Shams
Citations: 563
h-index: 13

AI 제어는 악의적인 목표를 가진 AI가 운영자의 의도를 왜곡하는 것을 방지하기 위한 기술들의 집합입니다. 일반적으로 AI 제어 연구에서는 단일 에이전트와 하나의 경로를 다루지만, 실제 환경에서는 여러 에이전트가 공유 인프라에서 작동하며, 가장 심각한 위험(모델 가중치 유출, 학습 과정 오염 등)은 여러 에이전트의 협력을 통해 발생할 가능성이 높습니다. 본 연구는 다중 에이전트 AI 제어에 대한 실증적 연구를 시작하고, 여러 에이전트가 공동으로 악의적인 목표를 달성하는 분산 공격을 형식화합니다. 우리는 FakeLab이라는 합성 AI 실험 환경 코드베이스(9개의 서비스, 86개의 정상 작업, 4가지 공격 목표)를 개발했습니다. 본 연구에서는 개별 에이전트 모니터링과 분산 공격 간의 관계를 평가하며, 에이전트 수, 협조 수준, 모델 능력 및 정확한 모니터링 구성을 다양하게 변화시킵니다. 중요한 결과는 '분산 효과'입니다. 공격에 참여하는 에이전트 수가 증가할수록, 개별 에이전트를 감시하는 시스템이 공격자를 탐지하기 어려워집니다. 둘째, 이는 정상 코드와 악성 코드의 비율에 의해 결정되는 것이 아니라, 모델 능력에 의존할 가능성이 높습니다. 셋째, 명시적인 계획 도구는 분산 효과를 증폭시키고, 공격 성공률을 최대 7배까지 증가시키며, 실행자의 미흡한 모니터링 이해도를 약간 개선합니다. 마지막으로, 취약한 '신뢰' 모니터는 대부분의 공격을 탐지하지 못하는 반면, 강력한 모니터는 탐지되지 않은 공격 성공률을 수십 배 이상 감소시킵니다. 본 연구에서 개발된 FakeLab과 제어 환경은 요청 시 안전 연구자들에게 제공될 예정입니다.

Original Abstract

AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studies a single agent in one trajectory, but real deployments run many agents over shared infrastructure, and the most severe risks (model-weight exfiltration, training-run poisoning) plausibly need several agents acting in concert. We initiate the empirical study of multi-agent AI control, formalising distributed attacks in which several agents jointly aim for a malicious goal. We develop FakeLab: a synthetic AI-lab codebase (9 services, 86 benign tasks, 4 attack objectives). We evaluate single agent monitoring against distributed attacks, varying the number of agents, their coordination, model capabilities and precise monitoring configuration. Our central finding is the fragmentation effect: as more agents coordinate to attack, per-agent monitoring becomes less likely to catch any of the attackers. Secondly, this is not driven by the ratio of benign to malicious code -- we argue this is likely dependent on model capability. Third, an explicit planner amplifies the fragmentation effect, raises the rate of completing the attack up to sevenfold, and mildly sharpens executors' poor monitor-theory-of-mind. Finally: a weak "trusted" monitor misses most attacks, while a stronger one cuts undetected success by more than an order of magnitude but remains imperfect. We are happy to share FakeLab, our control setting, with safety researchers on request.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!