2607.29190v1 Jul 31, 2026 cs.AI

CAGE: 타입화된 반환 불확실성 환경에서의 도구 사용 에이전트에 대한 인증된 권한 부여

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

Blaise Delattre
Blaise Delattre
Citations: 191
h-index: 5
Yang Cao
Yang Cao
Citations: 14
h-index: 2
Cong Wang
Cong Wang
Citations: 0
h-index: 0

도구를 사용하는 LLM 에이전트는 타입 정보가 있는 도구의 반환값을 활용하며, 이는 출처와 범주형 필드를 수치 값과 연결하는 기록으로 표현됩니다. 기존의 런타임 권한 게이트는 일반적으로 관찰된 반환값과 실행되는 동작을 허용하지만, 반환값이 소스에 어떻게 연결되었는지에 대한 작은 오류로부터 결정을 보호하지 못합니다. 본 연구에서는 특정 동작이 선언된 범위 내에서 유효한 상태를 유지하는지, 즉 정해진 오차 범위를 갖는 올바르게 연결된 반환값의 경우에도 유효한지를 질문합니다: 이는 하나의 허용 가능한 연결 오류와 경계가 설정된 수치 변화를 포함합니다. 우리는 범주형 및 수치 채널을 개별적으로 인증하는 것이 성립하지 않음을 증명했습니다. 각 채널에서 안전하다고 판단되는 작은 변화들이 결합되면 동일한 동작이 위험해질 수 있습니다. CAGE는 이러한 복합적인 영향을 직접적으로 인증하며, 이산적인 모든 가능성을 정확하게 나열하고, 각 분기 내의 연속적인 변화를 인증합니다. 합성 데이터, 정책-코드 기반 시스템, 규제 환경, 그리고 실제 거래 환경에서 CAGE는 기존 방식이 허용하는 잘못된 승인 사례를 줄이는 동시에, 여전히 상당수의 결정을 자율적으로 유지합니다. 정책이 실행 가능한 경우, CAGE-Exact는 해당 정책 자체를 인증하며, 그렇지 않은 경우에는 CAGE-Lip 및 CAGE-RS는 명시적이고 측정 가능한 신뢰도 가정을 바탕으로 학습된 게이트를 인증합니다.

Original Abstract

Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected against small errors in how the return was bound to its source. We ask whether a candidate action stays authorized over a declared neighborhood of plausible correctly bound returns: one admissible binding fault plus bounded numerical drift. We prove that certifying the categorical and numerical channels separately does not compose: perturbations that are safe on each channel alone can jointly turn the same action unsafe. CAGE certifies this joint neighborhood directly, enumerating the discrete branches exactly and certifying the continuous perturbation within each branch. Across synthetic, policy-as-code, regulatory, and real-transaction settings, CAGE removes the in-budget false allows that accurate pointwise gates admit, while keeping a useful fraction of decisions autonomous. When the policy is executable, CAGE-Exact certifies the policy itself; otherwise CAGE-Lip and CAGE-RS certify a learned gate under an explicit, measured fidelity assumption.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!