숨겨진 정보 사회 추론 게임에서의 신념 기반 LLM 에이전트 감사
Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games
숨겨진 정보가 있는 다중 에이전트 환경에서 LLM 에이전트를 평가하는 것은 어렵습니다. 최종 결과는 변동성이 크고, 에이전트가 왜 특정 결정을 내렸는지 보여주는 경우가 드물기 때문입니다. 본 연구에서는 9명의 플레이어가 참여하는 Werewolf 게임 환경을 통해 이를 연구합니다. 이 환경은 엄격한 코드 수준의 정보 격리를 적용하며, 외부 감사 프레임워크를 구축하여 숨겨진 역할에 대한 외부적인 신념 상태를 유지하고, 신념 업데이트 및 신념-행동 불일치를 구조화된 증거로 기록하며, 전략 변경 전에 문제 사례를 검토하는 방어적 오프라인 개선 루프를 지원합니다. 1,080개의 고정 게임을 통해 신념 비활성화, 활성 신념, 커널 제거, 캠프 제한, 소비 정책 및 높은 부하 조건 등을 비교했습니다 (A0/A1 시드 페어링 포함). 활성 신념 조건은 양호한 결과를 크게 향상시켰습니다. 200개의 시드 쌍을 사용한 A0/A1 비교에서 양호 측의 승률이 0.205에서 0.390으로 상승했습니다 (페어형 McNemar $χ^2 = 16.4$, p < 0.001), 또한 되돌릴 수 없는 마녀 독살 오류가 줄었습니다. 그러나 이러한 변화는 신념 내용 때문이라고 단정할 수 없습니다. 직접적인 행동-신념 일관성은 낮습니다 (약 0.21). 신념을 늑대인간에게만 부여하는 것이 양호 측에만 부여하는 것보다 더 효과적이라는 결과는 단순한 정보 획득 효과를 설명하지 못합니다. 따라서 본 연구에서는 이러한 효과를 연관성으로 보고, 그 메커니즘은 아직 해결되지 않았다고 판단합니다. 본 연구의 기여점은 감사 프레임워크 자체입니다. 이 프레임워크는 효과를 측정 가능하게 만들고, 낮은 직접적인 행동-신념 일관성을 드러내며, 신뢰할 수 없는 강제 소비 개입을 증거로 반박하고, 전략 효과와 부하 간의 혼동을 분리합니다. 따라서 외부 신념은 높은 노이즈의 숨겨진 정보 게임에서 주로 감사 가능한 인지 기반으로 작용하며, 동시에 의사 결정에 관련된 정보를 제공하여 불투명한 에이전트 행동을 재생 가능한 증거로 변환함으로써 더 안전하고 통제된 반복을 가능하게 합니다.
Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did. We study this in a 9-player Werewolf environment where agents act under strict, code-level information isolation, and we build an auditable framework that maintains an external belief state over hidden roles, logs belief updates and belief-action deviations as structured evidence, and supports a defensive offline improvement loop that reviews bad cases before any strategy change. Across 1,080 frozen games spanning belief-disabled, active-belief, kernel-ablation, camp-restricted, consumption-policy, and high-load arms, and including a seed-paired A0/A1 comparison, the active-belief condition is associated with substantially better good-side outcomes: in the 200-seed A0/A1 comparison the good-side win rate rises from 0.205 to 0.390 (paired McNemar $χ^2 = 16.4$, $p < 0.001$), with fewer irreversible witch-poison errors. We do not, however, attribute this shift to belief content. Direct action-belief consistency is low ($\approx 0.21$), and giving belief only to the werewolves helps the good side more than giving it only to the good side, which argues against a simple holder-benefit account; we therefore report the effect as an association and treat its mechanism as unresolved. The contribution is the audit framework itself: it makes the effect measurable, exposes low direct action-belief consistency, rejects an unreliable forced-consumption intervention with evidence, and separates strategy effects from load confounds. We accordingly position external belief in high-noise hidden-information games primarily as an auditable cognitive baseline that also carries decision-relevant signal, turning opaque agent behavior into replayable evidence for safer, controlled iteration.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.