검증 가능한 보안 에이전트 가이드레일
Provably Secure Agent Guardrail
대규모 언어 모델이 제한적인 생성 엔진에서 광범위한 실행 권한을 가진 에이전트로 진화하면서, 인공지능의 통제 불능은 인공지능 보안에 근본적인 위기를 초래합니다. 기존 방어 아키텍처는 주로 경험적인 의미 기반 가이드레일과 확률적 대규모 모델 판단 메커니즘에 의존하는데, 이러한 방식은 복잡한 의미 기반 기호 분리 공격에 직면했을 때 결정적인 보안 하한을 제공하지 못합니다. 본 논문에서는 이러한 경험적인 의미 기반 가이드레일의 한계를 극복하기 위해, 논리적 추론의 근본적인 제약을 바탕으로 에이전트의 새로운 보안 패러다임을 제시합니다. 이 패러다임을 바탕으로, 우리는 신경-기호 격리 아키텍처를 갖춘 실행 가능한 증거 기반 액션(ePCA) 프레임워크를 추가로 소개합니다. 이 프레임워크는 자연어에 대한 의미적 신뢰를 포기하고, 에이전트가 물리적인 작업을 수행하기 전에 의도를 손실 없이 1차 논리 수학적 제약 조건으로 형식화하도록 강제합니다. 거시적 및 미시적 2차원 동적 적대 시스템에 대한 경험적 평가 결과, 제안하는 형식 검증 메커니즘은 평가된 모든 시나리오에서 공격 성공률과 오탐율을 0%로 달성했으며, 매우 낮은 계산 지연 시간을 보였습니다. 본 연구는 명시적인 시스템 가정을 바탕으로 조건부 형식적 기반을 제공하며, 향후 지능형 시스템의 기본적인 방어 체계를 구축하기 위한 엔지니어링 패러다임을 제시합니다.
As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates a fundamental crisis in artificial intelligence security. Existing defense architectures heavily rely on empirical semantic guardrails and probabilistic large model adjudicators, mechanisms that fail to provide deterministic security lower bounds when facing complex semantic symbol decoupling attacks. To overcome this empirical semantic guardrail dilemma, this paper proposes a new security paradigm for agents based on the fundamental limitations of logical reasoning. Based on this paradigm, we further introduce an executable Proof-Constrained Action (ePCA) framework with a neural symbolic isolation architecture. This framework abandons semantic trust in natural language, forcing agents to losslessly formalize their intentions into first-order logical mathematical constraints before performing physical operations. Empirical evaluations of macroscopic and microscopic two-dimensional dynamic adversarial systems demonstrate that our formal verification mechanism achieves zero attack success rate and zero false positive rate across the evaluated scenarios, with extremely low computational latency. This research provides a conditional formal foundation under explicit system assumptions and an engineering paradigm for constructing the underlying defense foundation for future intelligent systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.