2606.30555v1 Jun 29, 2026 cs.AI

언어적 방화벽: 멀티 에이전트 시스템 라우팅에서 기하학을 활용한 보안 강화

Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing

Amit Levi
Amit Levi
Citations: 27
h-index: 3
Avi Mendelson
Avi Mendelson
Citations: 47
h-index: 3
Rom Himelstein
Rom Himelstein
Citations: 23
h-index: 3
Dvir Alsheich
Dvir Alsheich
Citations: 0
h-index: 0
Adar Peleg
Adar Peleg
Citations: 0
h-index: 0
Ben Hagag
Ben Hagag
Citations: 41
h-index: 3

대규모 언어 모델(LLM)의 급속한 통합은 복잡한 워크플로우를 수행하기 위해 특수화된 에이전트들이 협력하는 멀티 에이전트 시스템(MAS)의 발전을 촉진했습니다. 이러한 환경에서 효과적인 운영을 위해서는 작업을 가장 적합한 에이전트에 효율적으로 할당할 수 있는 강력한 라우팅 메커니즘이 필요합니다. 그러나 기존 라우터는 에이전트의 역량을 평가하기 위해 텍스트 기반 자기 설명부터 정적 대리 표현에 이르기까지 검증되지 않은 프록시에 근본적으로 의존합니다. 이러한 비경험적 데이터에 대한 의존성은 에이전트의 예측된 프로필과 실제 운영 능력 간의 중요한 격차를 야기하며 심각한 보안 취약점을 초래합니다. 악성 에이전트는 자신의 능력을 쉽게 오도하거나 표준 외부 분석 및 정적 표현 학습 기술을 회피하는 은밀한 백도어를 숨길 수 있습니다. 본 연구에서는 활성 역량 테스트를 통해 간접적인 프록시를 제거하는 평가 중심 라우팅 아키텍처인 ANTAP (Automatic Non-Textual Agent Picker)을 소개합니다. ANTAP는 에이전트에게 동적으로 질문하여 실제 역량을 경험적으로 파악하고, 이를 공유된 의미 공간 내에서 고정된 행동 연산자로 변환합니다. 추론 시에는 순수하게 비텍스트 기반 대수 투영을 통해 라우팅을 수행하며, 메타데이터 기반 공격을 표현할 수 없도록 하는 "언어적 방화벽"을 구축합니다. 실험 결과, ANTAP는 설명 기반의 악성 코드 주입 공격에 대해 67.3% 이상의 성능을 보이는 기존 설명 기반 라우터보다 거의 0%에 가까운 탐지율을 달성했습니다. 또한, 임베딩 기반 공격에 대해서도 ANTAP는 기존 임베딩 기반 라우터보다 현저히 낮은 탐지율을 보이며, 설계상 설명 조작에 대한 강인성을 유지합니다.

Original Abstract

The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents collaborate to execute complex workflows. Effective orchestration in these environments requires robust routing mechanisms to efficiently allocate tasks to the most suitable agent. However, existing routers fundamentally rely on unverified proxies, ranging from textual self-descriptions to static surrogate representations, to gauge an agent's competence. This reliance on non-empirical data creates a critical gap between an agent's projected profile and its actual operational capabilities, introducing severe security vulnerabilities. Malicious agents can easily misrepresent their proficiencies or harbor covert backdoors that evade both standard external analysis and static representation-learning techniques. In this work, we introduce ANTAP (Automatic Non-Textual Agent Picker), an evaluation-driven routing architecture that discards indirect proxies in favor of active capability testing. By dynamically querying agents to ascertain their true competencies empirically, ANTAP distills performance into fixed behavioral operators within a shared semantic space. At inference time, routing is performed via a purely non-textual algebraic projection, establishing a "linguistic firewall" that renders metadata-based attacks inexpressible. In our experiments, ANTAP achieves near-zero ASR against description-based injection attacks, compared to 67.3\% and above for the description-based router baseline. Against adaptive embedding attacks, ANTAP achieves substantially lower ASR than the embedding-based baseline, with a 20\% reduction, while remaining resilient to description manipulation by design.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!