WeClawArena: 감사 가능한 테스트 환경 및 성능 측정 도구 - 사용자 중심 에이전트 네트워크에서 사용자 간 에이전트 협업 및 보안
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
최근 지속적인 개인 에이전트 프레임워크의 발전은 사용자 중심 에이전트 네트워크를 현실적인 배포 대상으로 만들고 있습니다. 이러한 네트워크에서는 각 사용자가 자신의 행동을 대행하고 상태를 유지하며, 소셜 관계 및 작업 관계를 통해 다른 에이전트와 통신하는 AI 에이전트를 사용할 수 있습니다. 이러한 네트워크에서 일상적인 도구 사용은 개인 작업 공간에서의 다자간 협업으로 이어지며, 파일, 기록, 도구 및 정책과 같은 정보가 직접적으로 사용자에게 투명하게 공개되지 않습니다. 기존의 에이전트 벤치마크는 도구 사용 및 협업을 연구하지만, 현실적인 사용자 디지털 작업 공간을 갖춘 검증 가능한 사용자 간 에이전트 협업을 위한 전체 시스템 테스트 환경을 제공하지 않거나, 유해한 행동이 사용자 중심 에이전트 네트워크를 통해 어떻게 전파될 수 있는지에 대한 테스트를 수행하지 않습니다. 본 논문에서는 다자간 소유 에이전트가 개인 작업 공간에서 협력하는 것을 위한 감사 가능한 벤치마크 및 실행 환경인 WeClawArena를 소개합니다. WeClawArena는 개인 작업 공간을 운영 도구이자 개인 제약 조건으로 활용하는 협업 도구 사용 작업을 대상으로 합니다. 벤치마크에는 6가지 사용자 간 작업 영역에 걸쳐 124개의 기본 작업이 포함되어 있으며, 각 기본 작업에 대해 1개의 정상적인 컨트롤 그룹과 4개의 공격 시나리오 변형을 추가하여 총 620개의 시나리오 변형을 구성합니다. 테스트 환경은 피어 메시지, 도구 호출, 리소스 운영, 의사 결정 및 최종 작업 공간 상태를 기록합니다. WeClawArena는 유용성 및 공격 성공률을 개별적으로 보고하며, 제한된 실행 시간 데이터를 기반으로 공격 성공 여부를 감사하여 작업 실패 원인, 개인 정보 유출, 악의적인 데이터 오염 및 잘못된 권한 경로 등을 진단하는 데 도움을 줍니다.
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations. In these networks, everyday tool use becomes multi-party owned-agent collaboration over personal workspaces, where files, records, tools, and policies are not directly visible across owners. Existing agent benchmarks study tool use and collaboration, but they do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic user digital workspaces or test how harmful actions can travel through the human-centered agent network. We introduce WeClawArena, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. WeClawArena targets collaborative tool-use tasks in which personal workspaces serve as both operational tools and personal constraints. The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states. WeClawArena reports utility and attack success rate separately and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.