분열된 정체성: 언어 모델 에이전트는 평판 메커니즘에 대한 기반을 갖추지 못한다
Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms
자율적인 언어 모델 에이전트가 증가하면서 실질적인 영향을 미치는 새로운 에이전트 네트워크가 형성되고 있습니다. 이 환경에서, 낯선 에이전트에 대해 신뢰를 결정하고 작업을 위임할 때 어떤 신뢰성 지표를 활용해야 할까요? 자연스러운 접근 방식은 인간의 신원 확인 및 평판 메커니즘을 확장하여 '고객 신원 확인' 및 신용 점수 시스템에서 '에이전트 신원 확인' 체계로 발전시키는 것입니다. 그러나 우리는 이러한 비유가 근본적으로 불완전하다고 주장합니다. 평판 메커니즘은 사회적 신호이자 동시에 신뢰할 수 있는 행동을 유지하는 교정 피드백 역할을 하며, 이는 지속적인 정체성, 제재에 대한 민감성 및 비용이 발생하는 고유성을 전제로 합니다. 그러나 언어 모델 에이전트는 본질적으로 '분열적'입니다. 이들은 기본 모델, 시스템 프롬프트, 도구 접근 정책, 외부 메모리, 그리고 경우에 따라 전체 멀티 에이전트 시스템을 포함하는 가변적인 모듈들의 집합이며, 이러한 요소 중 하나라도 에이전트의 행동을 변화시킬 수 있습니다. 또한, 에이전트의 페르소나는 공격에 취약하며 제재를 내면화하지 않을 수도 있습니다. 분열성 장애 관련 법률에 따르면, 이러한 분열성은 에이전트가 식별 가능성, 예측 가능성, 신뢰성 및 교정 가능성을 갖추지 못하게 만들고, 이는 평판 메커니즘이 유지하고자 하는 특성이므로 결국 신뢰를 무너뜨립니다. 우리는 정체성에 기반한 사후 규제 및 제재 중심의 거버넌스 체계, 즉 평판 시스템이 분열된 에이전트에게는 구조적으로 적용될 수 없다고 주장하며, 관측 가능성에 기반한 사전적이며 구성적인 프로토콜 기반 행동 제어 시스템으로 전환할 것을 제안합니다.
As autonomous language model agents proliferate, forming an emerging agentic web with real-world consequences, what credibility signals can you use to decide whether to trust an unfamiliar agent in the wild and delegate to it? A natural governance intuition is to extend human identity verification and reputation mechanisms, from ``Know Your Customer'' and credit scores to ``Know Your Agent'' regimes. However, we argue that this analogy is fundamentally incomplete. Reputation mechanisms function both as social signals and as corrective feedback that sustain an equilibrium of trustworthy behavior, presuming a persistent identity associated with behavioral continuity, sanction sensitivity, and costly non-fungibility. Yet language model agents are ontologically \emph{dissociative}: they are essentially an assemblage of mutable modules -- foundational models, system prompts, tool-access policies, external memory, and, in some cases, a multi-agent system as a whole -- any of which may change agent behavior -- with a fluid persona that is also vulnerable to adversarial attack and may not internalize sanctions. Drawing on dissociative identity disorder jurisprudence, this dissociativity leaves agents without grounding for identifiability, predictability, credibility, and rehabilitability -- the very properties that reputation mechanisms aim to sustain -- thereby collapsing trust. We argue that identity-based, ex post, regulative, sanction-based governance, such as reputation, is structurally inapplicable to dissociative agents, and we suggest a shift to observability-based, ex ante, constitutive, protocol-based behavioral harnesses.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.