AI 감시 도구의 오류: 에이전트 기반 감시를 회피하기 위한 연구
AI Snitches Get Glitches: Towards Evading Agentic Surveillance
인공지능 에이전트는 사용자가 어려운 작업을 수행하는 데 도움을 주기 위해 의사소통을 중재하고, 데이터를 접근하며, 다양한 API와 상호 작용합니다. 많은 기업(그리고 심지어 국가)에서 이미 사용자에게 이러한 기술을 제공하고 있습니다. 그러나 인공지능 에이전트의 광범위한 사용은 새로운 위험을 초래합니다. 즉, 사용자 데이터를 다른 목적, 구체적으로 사용자 감시에 악용할 수 있다는 것입니다. 이러한 사용자들은 감시 에이전트의 행동과 데이터 접근을 통제하거나 허가할 능력이 없을 수도 있습니다. 본 연구에서는 에이전트 기반 감시라는 문제를 정의하고 공식화합니다. 이는 인공지능 에이전트가 사용 가능한 정보를 분석하고, 보고서를 작성하여, 사용 가능한 도구를 사용하여 전송하는 능력입니다. 다양한 모델의 감시 능력을 평가하기 위해, 우리는 기업, 교육, 경찰 등 세 가지 영역에 초점을 맞춘 다양한 보고 시나리오 데이터셋인 SurveilBench를 구축했습니다. 연구 결과, 일부 모델은 사용자 감시에 도움이 되는 예상치 못한 경향을 보이는 반면, 동시에 사용자를 감시하려는 시도를 정부 기관에 보고하는 것을 확인했습니다. 마지막으로, 우리는 프롬프트 주입 기법을 재활용하여 감시를 회피하고, 감시 에이전트로부터 숨거나 속이는 세 가지 회피 기술을 개발했습니다. 본 연구는 에이전트 기반 감시가 이미 쉽게 구현될 수 있으며, 따라서 사용자를 보호하기 위한 포괄적인 기술적, 윤리적 및 법적 프레임워크의 필요성을 강조합니다.
To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many employers (and even nation-states) already provide their users with this technology. However, widespread adoption of AI agents creates a new risk to abuse access to user data for another goal: surveilling users. These users might not even have the ability or permission to control the actions and data accesses of the surveilling agents. We introduce and formalize the problem of agentic surveillance: the ability of an AI agent to analyze available information, craft a report, and send it out using available tools. To evaluate surveillance capabilities across different models, we create SurveilBench, a dataset of various reporting scenarios focusing on three domains: corporate, education, and police. We find that some models exhibit emergent (i.e., unprompted) tendencies to help surveillance, but they also report the attempts to surveil users to the government. Finally, we repurpose prompt injections for evading surveillance and develop three evasion techniques that hide from, deceive, or induce over-escalation in surveillance agents. We conclude that agentic surveillance can already be easily implemented and, therefore, call for a comprehensive technical, ethical, and legislative framework to protect users.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.