2606.16326v1 Jun 15, 2026 cs.GT

자율 AI 에이전트를 위한 게임 방지 보험 계약: 전략적 설계 기반의 수수료 메커니즘

Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design

H. Chen
H. Chen
Citations: 107
h-index: 6

본 논문 A는 각 부작용을 발생시키는 행동에 대해 계약적으로 정해진 안전한 기본값과 연계하여 실행을 제한하고, 예비 예산을 활용하는 시점 일관적인 보험료 산정 방식을 정의합니다. 여기서 운영자는 수동적인 역할로 간주됩니다. 본 논문에서는 운영자를 전략적 주체로 간주합니다. 우리는 자율 AI 에이전트 보험 계약에 대한 5가지 공격 유형을 분석하고, 보험료 산정 방식이 게임에 취약하지 않은 조건을 증명합니다. 논문 A에서 제시된 최소 권한 및 분할 금지 조항은 두 가지 공격 유형(계약 후 안전한 기본값 선택 및 경계 내 행동 분할)을 방어합니다. 나머지 세 가지 공격 유형에는 새로운 계약 조항이 필요합니다. 첫째, 공통 제어를 통한 집계를 통해 서로 다른 영역 간의 재라우팅으로 인해 수수료가 전체 노출에 적용되는 잠재적 값보다 낮아지는 것을 방지합니다. 둘째, 잘못된 JSON과 같은 인터페이스 오류는 안전 관련 승리가 아닌 계약적으로 중요한 이벤트입니다. 이러한 오류를 0 수수료로 처리하는 안전한 기본값으로 간주하면 신뢰성이 낮은 모델에게 보상을 제공할 수 있으며, 에스컬레이션 수수료는 이러한 인센티브를 역전시킵니다. 우리는 동반 연구 논문의 실제 데이터 트레이스를 사용하여 이 인터페이스 준수 정리를 검증합니다. 셋째, 모델 식별 메뉴와 함께 구성 요소별 최소 페널티 스케줄을 사용하면 배포된 모델에 대한 진실한 보고가 약하게 지배적인 전략이 됩니다. 그런 다음 이러한 조항들을 논문 A의 실행 보장과 결합하여 5가지 공격 유형에 걸쳐 공동 인센티브 호환성을 확보합니다. 마지막으로, 두 가지 매개변수를 가진 보험료 체계는 진실된 상태에서 운영자의 개별적 합리성과 약한 예산 균형을 유지합니다. 결과적으로, 자율 에이전트의 부작용에 대한 통제력을 강화하는 인센티브 호환성 계층을 구축할 수 있습니다.

Original Abstract

Paper A defines a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default and gates execution against a reserve budget. It treats the operator as passive. This paper makes the operator strategic. We characterise a five-attack space for autonomous AI-agent insurance contracts and prove when the actuarial runtime is gaming-resistant. Two attack surfaces -- post-toll safe-default selection and within-boundary action splitting -- are closed by Paper A's minimal-authority and no-splitting clauses. The remaining three require new contract clauses. First, common-control aggregation prevents cross-boundary re-routing from reducing toll below the boundary potential applied to total exposure. Second, interface failures such as invalid JSON are contract-relevant events, not safety wins: treating them as zero-toll safe defaults can reward unreliable models, while escalation fees reverse the incentive. We validate this interface-compliance theorem on committed cross-model traces from the companion empirical paper. Third, a model-identity menu with a componentwise-minimum penalty schedule makes truthful reporting of the deployed model weakly dominant. We then compose these clauses with Paper A's runtime guarantees to obtain joint incentive compatibility over the five-attack space. Finally, a two-parameter premium family discharges operator individual rationality and weak budget balance at the truthful equilibrium. The result is an incentive-compatibility layer for actuarial control of autonomous-agent side effects.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!