2606.05793v1 Jun 04, 2026 cs.CL

CollabBench: 다양한 사용자와의 능동적인 상호작용을 통한 LLM의 협업 능력 평가 및 향상

CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement

Aimin Zhou
Aimin Zhou
Citations: 5
h-index: 1
Xiangfeng Wang
Xiangfeng Wang
Citations: 43
h-index: 2
Yuanhao Liu
Yuanhao Liu
Citations: 17
h-index: 3
Liang Dou
Liang Dou
Citations: 40
h-index: 2
Hong Qian
Hong Qian
Citations: 179
h-index: 9
Zihan Zhou
Zihan Zhou
Citations: 550
h-index: 5
Jingwen Yang
Jingwen Yang
Citations: 119
h-index: 4
Hanjie Ge
Hanjie Ge
Citations: 1
h-index: 1
Haotian Shi
Haotian Shi
Citations: 10
h-index: 1
Zongbao Zhang
Zongbao Zhang
Citations: 0
h-index: 0

LLM 기반 에이전트는 개별 작업에서 뛰어난 성능을 보이지만, 실제 인간 파트너와의 효과적인 협력은 여전히 어려운 과제입니다. 기존의 대부분의 대화 수준 협력 연구는 실질적인 상호작용과 행동 실행에 대한 고려가 부족하며, 이는 맥락화된 몰입형 협력을 가능하게 하는 협력 게임 환경의 필요성을 야기합니다. 이에 본 논문에서는 협력 게임에서 협업 에이전트를 평가하고 훈련하기 위한 벤치마크인 CollabBench를 제안합니다. CollabBench는 다양한 플레이어의 행동을 모델링하는 '다양한 플레이어 프로필 시뮬레이션' 파이프라인과, 추론, 의사소통 및 액션을 에이전트 기반 실행을 통해 통합하고, 작업 효율성과 정서적 적응을 균형 있게 최적화하는 '협력 에이전트 훈련 패러다임'을 특징으로 합니다. 또한, CollabBench는 다양한 성격 하에서 체계적인 평가를 위해 기존 환경을 CWAH-MultiPlayer 및 Cook-MultiPlayer로 확장했습니다. 효율성과 정서 관련 지표에 대한 실험 결과, 저희가 훈련한 모델은 기본 모델보다 19.5% 더 높은 효율성과 24.4% 향상된 정서적 성능을 달성했습니다. 추가 분석을 통해 기존 모델의 주요 협력적 한계를 밝히고, 향후 협력 훈련에 대한 통찰력을 제공합니다.

Original Abstract

While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conversation-level collaborative studies lack grounded interaction and behavioral execution, motivating the need for cooperative game environments that enable contextualized and immersive collaboration. To this end, this paper proposes CollabBench, a benchmark for evaluating and training collaborative agents in cooperative games. CollabBench features a Diverse Player Profile Simulation pipeline to model varied players behaviors, and a Collaborative Agentic Training paradigm that unifies reasoning, communication, and action via agentic rollouts, optimized with a hybrid reward balancing task efficiency and affective adaptation. We further extend classic environments to CWAH-MultiPlayer and Cook-MultiPlayer for systematic evaluation under diverse personalities. Experiments with efficiency and affective metrics show that our trained models outperform base models, achieving 19.5% higher efficiency and 24.4% improved affective performance. Further analysis reveals key collaborative limitations of existing models and offers insights for future collaborative training.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!