믿을 수 없지만 부패하다: 다중 에이전트 거버넌스 시스템에서의 부패 평가
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
대규모 언어 모델(LLM)은 고위험 공공 업무 프로세스를 위한 자율 에이전트로 점점 더 많이 제안되고 있지만, 권한이 부여되었을 때 이러한 모델이 기관 규칙을 준수하는지에 대한 체계적인 증거는 부족합니다. 본 연구는 기관 AI의 무결성이 배포 후의 가정이라기보다는 배포 전의 필수 요건으로 간주되어야 한다는 증거를 제시합니다. 다양한 권한 구조 하에서 에이전트가 공식적인 정부 역할을 수행하는 다중 에이전트 거버넌스 시뮬레이션을 평가하고, 28,112개의 트랜스크립트 세그먼트에 대해 독립적인 기준에 기반한 평가를 통해 규칙 위반 및 남용 결과를 측정했습니다. 본 연구는 이러한 주장을 발전시키는 동시에, 핵심적인 기여는 경험적인 증거를 제공하는 것입니다. 포화 상태에 도달하지 않은 모델의 경우, 모델 자체의 특성보다 거버넌스 구조가 부패 관련 결과에 더 큰 영향을 미치며, 체제 및 모델-거버넌스 조합에 따라 상당한 차이가 나타납니다. 경량화된 안전 장치는 일부 환경에서 위험을 줄일 수 있지만, 심각한 실패를 지속적으로 방지하지는 못합니다. 이러한 결과는 기관 설계가 안전한 위임의 필수 조건임을 시사합니다. LLM 에이전트에 실제 권한을 부여하기 전에, 시스템은 시행 가능한 규칙, 감사 가능한 로그, 그리고 중요한 작업에 대한 인간의 감독을 갖춘, 거버넌스와 유사한 제약 조건 하에서 스트레스 테스트를 거쳐야 합니다.
Large language models are increasingly proposed as autonomous agents for high-stakes public workflows, yet we lack systematic evidence about whether they would follow institutional rules when granted authority. We present evidence that integrity in institutional AI should be treated as a pre-deployment requirement rather than a post-deployment assumption. We evaluate multi-agent governance simulations in which agents occupy formal governmental roles under different authority structures, and we score rule-breaking and abuse outcomes with an independent rubric-based judge across 28,112 transcript segments. While we advance this position, the core contribution is empirical: among models operating below saturation, governance structure is a stronger driver of corruption-related outcomes than model identity, with large differences across regimes and model--governance pairings. Lightweight safeguards can reduce risk in some settings but do not consistently prevent severe failures. These results imply that institutional design is a precondition for safe delegation: before real authority is assigned to LLM agents, systems should undergo stress testing under governance-like constraints with enforceable rules, auditable logs, and human oversight on high-impact actions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.