2604.22879v1 Apr 24, 2026 cs.MA

단일 에이전트 정렬을 넘어: 다중 에이전트 시스템에서 발생하는 컨텍스트 단편화 위협 방지

Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems

Ming Gong
Ming Gong
Citations: 28
h-index: 2
Jie Wu
Jie Wu
Citations: 81
h-index: 5

본 연구에서는 새로운 보안 위협인 '컨텍스트 단편화 위협(Context-Fragmented Violations, CFV)'을 식별하고 정의합니다. CFV는 개별 에이전트의 행동은 현지적으로 안전하고 합리적으로 보이지만, 중요한 정책 정보가 서로 다른 부서의 폐쇄된 컨텍스트에 분산되어 있어 전체적으로 조직 정책을 위반하는 경우를 의미합니다. 기존의 프롬프트 기반 정렬 메커니즘과 통합형 인터셉터는 이러한 컨텍스트 경계를 넘나드는 위협에 효과적으로 대응하지 못합니다. 우리는 분산형 제로 트러스트 강제 아키텍처인 '분산 센티넬(Distributed Sentinel)'을 제안하며, 이 아키텍처는 '의미 기반 추적 토큰(Semantic Taint Token, STT) 프로토콜'을 도입합니다. 경량화된 사이더(sidecar) 프록시를 통해, 당사 시스템은 원시 데이터를 노출하지 않고도 조직 경계를 넘어 보안 상태를 전파하며, 이를 통해 다양한 도메인 간 정책 검증을 위한 가상 그래프 시뮬레이션을 가능하게 합니다. 우리는 9가지 유형의 현실적인 에이전트 간 위반 시나리오를 포함하는 종합적인 벤치마크인 '팬텀 에코시스템(PhantomEcosystem)'을 구축했습니다. 이 벤치마크에서 분산 센티넬은 F1 점수가 0.95로, 프롬프트 기반 필터링의 0.85, 규칙 기반 DLP의 0.65보다 우수한 성능을 보였으며, 전체 지연 시간은 106ms (A100에서 16ms의 검증 및 90ms의 엔티티 추출)입니다. 또한, 외부 강제 메커니즘의 필요성을 실증적으로 검증하기 위해, 에이전트별 도메인 세계 모델을 사용하는 실행 중심의 다중 에이전트 워크플로우에서 8개의 최첨단 LLM을 평가했습니다. 모든 모델에서 상당한 위반율(14-98%)이 관찰되었으며, 특히 다른 도메인 간 데이터 흐름에서 위반율이 더 높게 나타났습니다. 이러한 결과는 자기 회피(self-avoidance)가 신뢰할 수 없으며, 다중 에이전트 시스템의 보안은 개별 에이전트 위에 작동하는 중앙 집중식 강제 계층으로부터 이점을 얻을 수 있음을 시사합니다.

Original Abstract

We identify and formalize a novel security risk: Context-Fragmented Violations (CFVs) - a class of policy breaches where individual agent actions appear locally safe and reasonable, yet collectively violate organizational policies because critical policy facts are siloed in different departments private contexts. Existing prompt-based alignment mechanisms and monolithic interceptors are poorly matched to violations that span contextual islands. We propose Distributed Sentinel, a distributed zero-trust enforcement architecture that introduces the Semantic Taint Token (STT) Protocol. Through lightweight sidecar proxies, our system propagates security state across organizational boundaries without exposing raw cross-domain data, enabling Counterfactual Graph Simulation for cross-domain policy verification. We construct PhantomEcosystem, a comprehensive benchmark comprising 9 categories of realistic cross-agent violation scenarios with adversarially balanced safe controls. On this benchmark, Distributed Sentinel achieves F1 = 0.95 with 106ms end-to-end latency (16ms verification + 90ms entity extraction on A100), compared to 0.85 F1 for prompt-based filtering and 0.65 for rule-based DLP. To empirically validate the need for external enforcement, we evaluate eight frontier LLMs in execution-oriented multi-agent workflows with per-agent domain world models. All models exhibit substantial violation rates (14-98%), with cross-domain data flows showing systematically higher violation rates than same-domain flows. These results indicate that self-avoidance is unreliable and that multi-agent security benefits from a centralized enforcement layer operating above individual agents.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!