MAStrike: 셰이플리 값을 활용한 다중 에이전트 시스템의 공모형 레드 팀 공격
MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems
계층적 다중 에이전트 시스템(MAS)은 금융 및 소프트웨어 엔지니어링 등 다양한 분야에서 중요한 업무에 빠르게 도입되고 있습니다. 이러한 시스템에서 안전과 보안은 역할별로 특화된 에이전트에 분산되어 있으며, 이는 권한 상승 및 에이전트 간의 공모와 같은 조정된 적대적 행동 하에서 공격 표면을 크게 확대합니다. 기존의 MAS 레드 팀 접근 방식은 여전히 제한적입니다. 이러한 방식들은 휴리스틱 기반으로 대상 에이전트를 선택하고 격리된 메시지 스트림을 조작하며, 시스템 안전에 가장 큰 영향을 미치는 에이전트가 무엇인지, 그리고 손상된 에이전트가 방어를 우회하기 위해 어떻게 협력할 수 있는지와 같은 중요한 질문에 대한 답을 제공하지 못합니다. 본 논문에서는 계층적 MAS에서 공모형 레드 팀 공격을 위한 폐루프 프레임워크인 MAStrike를 제안합니다. 우리는 첫 번째로, MAS의 에이전트 수준에서의 셰이플리 값 분석을 제안하여 각 에이전트가 특정 작업 분포 하에서 시스템 견고성에 기여하는 정도를 정량화합니다. 이러한 기여도 기반으로 MAStrike는 취약한 에이전트 연합을 식별하고 조정된, 역할 인지적인 적대적 조작을 생성합니다. 이러한 공격은 구조화된 원인 진단을 통해 반복적으로 개선되며, 손상되지 않은 에이전트가 적대적 시도를 차단하는 실패 사례를 분석하여 책임 소재를 파악합니다. 또한, 금융, 소프트웨어 엔지니어링 및 CRM 등 다양한 계층적 토폴로지와 도메인을 포괄하는 종합적인 MAS 레드 팀 벤치마크와 제어 가능한 환경을 구축했습니다. 여러 최첨단 모델을 기반으로 구축된 MAS에 대한 광범위한 실험 결과, MAStrike가 휴리스틱 기반의 기존 방법보다 훨씬 우수한 성능을 보입니다. 또한, 분석 결과에서 셰이플리 값 분포 및 에이전트 간의 고차원적 상호 작용 구조를 밝혀내어, 기존의 단일 에이전트 또는 템플릿 기반 방법으로는 간과될 수 있는 중요한 취약점과 협력 패턴을 드러냅니다.
Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering. In these systems, safety and security are inherently distributed across role-specialized agents, significantly expanding the attack surface, particularly under coordinated adversarial behaviors such as privilege escalation and cross-agent collusion. Existing red-teaming approaches for MAS remain limited: they rely on heuristic selection of target agents and perturb isolated message streams, leaving critical questions unanswered as which agents are most responsible for system safety, and how compromised agents can coordinate to bypass defenses. We propose MAStrike, a closed-loop framework for collusive red-teaming in hierarchical MAS. We propose the first agent-level Shapley value analysis for MAS, quantifying each agent's marginal contribution to system robustness under task-specific distributions. GGuided by this attribution, MAStrike identifies vulnerable agent coalitions and generates coordinated, role-aware adversarial manipulations. These attacks are iteratively refined through structured causal diagnosis, attributing failure cases to uncompromised agents that block adversarial attempts. We further build a comprehensive MAS red-teaming benchmark and controllable environments spanning diverse hierarchical topologies and domains, including finance, software engineering, and CRM. Extensive experiments across MAS built on multiple frontier models show that MAStrike substantially outperforms heuristic baselines. Our analysis further uncovers non-trivial Shapley value distributions and higher-order interaction structures among agents, revealing critical vulnerabilities and coordination patterns that are overlooked by prior single-agent or template-based methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.