능력은 있지만 부주의한: 컴퓨터 사용 에이전트는 문맥적 무결성을 준수하는가?
Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
컴퓨터 사용 에이전트(CUA)는 현재 이메일, 캘린더, 할 일 목록과 같은 개인 애플리케이션에서 사용자를 대신하여 작동합니다. 이러한 애플리케이션 간의 접근성은 유용하지만, 프라이버시 위험을 야기하며 이는 대부분 간과되어 왔습니다. 에이전트가 한 문맥에서 작업할 때, 해당 문맥에 부적절한 정보를 다른 곳에서 가져올 수 있습니다. 따라서 우리는 이러한 위험을 실행 가능한 형태로 만들고, 결정적으로 점수 매김이 가능한 시나리오를 제공하는 평가 도구인 AgentCIBench를 소개합니다. 우리는 CUA에서 발생하는 세 가지 일반적인 오류 방식을 대상으로 합니다: UI에서 작업 대상 옆에 위치한 금지된 항목을 가져오는 '시각적 근접성', 과도하게 상세한 개인 정보를 요청하지 않은 프롬프트에 대해 제공하는 '작업 모호성 과잉 공유', 그리고 콘텐츠를 적절하지 않은 수신자에게 보내는 '수신자 불일치'. 우리는 15개의 최첨단 에이전트를 평가한 결과, 놀라울 정도로 높은 오류율을 발견했습니다. 15개 중 11개가 시나리오의 50% 이상에서 정보를 유출했으며, 평균 유출률은 67.9%였습니다. 또한, 에이전트가 작업을 완료하기 위해 환경 내에서 전체적으로 작동하는 경우에도 동일한 오류가 지속되었습니다. 우리는 AgentCIBench를 공개하여 더 안전한 컴퓨터 사용 에이전트 개발을 장려하고, 문맥 정보 공개 테스트를 배포 전 안전 점검으로 자리매김하고자 합니다.
Computer-use agents (CUAs) now act on a user's behalf across personal applications such as email, calendars, and to-do lists. This cross-application access is useful, but it also creates a privacy risk that has been largely overlooked: when an agent works in one context, it can pull in information from another that is inappropriate in that context. Hence, we introduce AgentCIBench, an evaluation harness that turns this risk into executable, deterministically scored scenarios. We target three common failure modes in CUAs: visual co-location, where the agent pulls in prohibited items that sit next to the task target in the UI; task-ambiguity overshare, where the agent dumps dense personal state in response to an under-specified prompt; and recipient misalignment, where the agent sends content to an addressee for whom it is inappropriate. We evaluate 15 frontier agents and find a surprisingly high failure rate: 11 of 15 leak on more than 50% of scenarios, with an average leakage of 67.9%, and the same failures persist when agents act end-to-end in the environment to complete the task. We release AgentCIBench to encourage the development of safer computer-use agents and position contextual disclosure testing as a pre-deployment safety check.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.