에이전트 자동화가 수익성이 높아지는 시점: 추적 경제 기반 인수 심사를 통한 자율 AI 위험의 정량화 및 보험 적용
When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting
AI 에이전트는 이제 운영 시스템에서 되돌릴 수 없는 작업을 수행할 수 있지만, 에이전트가 초래하는 손실은 여전히 명확하게 책정되지 않거나 가격 결정되지 않고 이전되지 않습니다. 서비스 제공업체는 종종 결과적 손해에 대한 책임을 부인하고, 사용자들은 보상받지 못하는 손실을 입으며, 기존의 인간 검토 방식은 자동화로 인한 효율성 향상을 제한합니다. 본 연구에서는 자율 AI 배포가 실패 위험에도 불구하고 경제적으로 타당해지는 시점에 대해 질문합니다. 저희는 고객-작업-추적 에피소드 수준에서 위험을 정량화하고, 이를 보험을 통해 이전함으로써 해결책을 제시합니다. 자동화는 예상되는 이점이 보험료, 제어 비용 및 잔여 위험을 초과할 때 허용 가능합니다. 이를 위해서는 명확하게 정의된 역할과 제한적인 권한, 그리고 비교 가능한 추적이 필요합니다. 본 연구에서는 도구 사용 추적을 고객 노출 및 청구 가능한 손실로 매핑하는 '추적 경제 기반 인수 심사' 방법을 제시하며, 이를 가격 결정, 제어 및 위험 이전에 활용합니다. 이 방법은 LLM 판사가 아닌, 결정론적인 경제적 라벨을 사용합니다. 저희가 개발한 추적-손실 테스트 환경에서, 추적 경제 기반 가격 결정 방식은 평균 절대 오차(MAE)를 17,700달러에서 569달러로 줄이고, 불공정한 교차 보조금을 제거했습니다. 전문가 검토 결과, 300개의 추적 데이터에 대해 295개의 라벨이 변경 없이 그대로 유지되었습니다. 또한, 실제 소프트웨어 엔지니어링(SWE) 추적 데이터를 기반으로, 추적 정보에 따른 제어를 통해 CVaR95를 72% 감소시켰습니다. Theorem~1은 유한 샘플 범위 조건을 제시합니다. 저희는 코드, 라벨 및 감사 자료를 공개합니다.
AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or transferred. Providers often disclaim consequential damages, users are left with uncompensated losses, and default human review limits the efficiency gains of automation. We ask when autonomous AI deployment can become economically acceptable despite failure risk. Our answer is to quantify risk at the customer-task-trace episode level and transfer it through insurance. Automation is acceptable when its expected benefit exceeds the premium, control cost, and remaining risk. This requires a defined role with bounded permissions and comparable traces. We introduce trace-economic underwriting, which maps tool-use traces to customer exposure and claimable loss, then uses this representation for pricing, control, and risk transfer. It uses deterministic economic labels rather than an LLM judge. In our trace-to-loss testbed, trace-economic pricing reduces pricing MAE from $17.7K to $569 and removes regressive cross-subsidy. A 300-trace expert audit accepts 295 labels unchanged. On 1,000 real SWE-smith traces, trace-conditioned controls reduce CVaR95 by 72%. Theorem~1 gives a finite-sample scope condition. We release code, labels, and audit sheets.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.