CaM-Wolf: 인과 관계 인식 능력을 갖춘 다중 모드 에이전트 - 사회 추론 게임을 위한 연구
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
사회 추론 게임(SDG), 예를 들어 Werewolf은 AI 에이전트를 테스트하는 데 유용한 플랫폼으로 자리 잡았습니다. 이러한 게임은 추론, 기만 및 협력과 같은 복잡한 사회적 기술을 요구합니다. 최근 대규모 언어 모델(LLM)의 발전은 SDG 에이전트 개발에 상당한 진전을 가져왔지만, 현재 대부분의 접근 방식은 텍스트 기반이며 인간 사회 상호 작용의 근본적인 다중 모드 특성을 간과하고 있습니다. 이러한 격차를 해소하기 위해, 우리는 CaM-Wolf을 소개합니다. CaM-Wolf은 다중 모드 인식 및 생성을 통합하는 최초의 SDG 에이전트입니다. CaM-Wolf은 다른 플레이어로부터 비디오 입력을 처리하고, 강화 학습을 통해 훈련된 인과 관계 인식 추론기를 사용하여 관찰 가능한 행동과 숨겨진 역할 간의 논리적 연결을 확립하며, 애니메이션 아바타를 통해 자신을 표현합니다. 우리의 실험 및 사용자 연구 결과는 CaM-Wolf이 우수한 에이전트 게임 성능을 달성하고 인간-AI 상호 작용의 품질을 향상시킨다는 것을 보여줍니다. 이 연구는 미묘한 사회적 역학에 참여할 수 있는 더욱 인간과 유사한 AI 에이전트를 개발하는 데 중요한 진전을 나타냅니다. 저희 코드 및 관련 자료는 다음 링크에서 확인할 수 있습니다: https://3dagentworld.github.io/avatar_wolf.
Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that is fundamental to human social interaction. To bridge this gap, we introduce CaM-Wolf, the first SDG agent that integrates multimodal perception and generation. CaM-Wolf processes video inputs from other players, employs a causal-aware Reasoner trained via reinforcement learning to establish logical chains between observable behaviors and hidden roles, and presents itself through an animated avatar. Our experiments and user study show that CaM-Wolf achieves superior agent gameplay performance and enhances the quality of human-AI interaction. This work represents a significant advancement towards creating more human-like AI agents capable of participating in nuanced social dynamics. Our code is available at https://3dagentworld.github.io/avatar_wolf.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.