ALMANAC: 에이전트 협업을 위한 행동 수준의 정신 모델 주석 데이터셋 – 인간-인간 협업 사례
Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration
최근 LLM(Large Language Model) 기반 에이전트 기술은 다단계 추론, 계획 수립 및 도구 활용과 같은 복잡한 인지 능력을 가능하게 하여, 이러한 에이전트들이 점차 인간의 협력 파트너로서 자리 잡고 있습니다. 효과적인 협업을 위해서는 참여자들이 자신의 사고 과정, 파트너의 의도, 그리고 공동 목표에 대한 정신 모델을 지속적으로 유지하고 일치시켜야 합니다. 그러나 현재 대부분의 에이전트는 주로 작업 완료에 최적화되어 있으며, 행동 수준의 정신 모델 주석이 포함된 실제 인간 협업 데이터가 부족하여 에이전트들이 프로세스 수준의 협업 역량을 갖추도록 유도하기 어렵습니다. 이러한 간극을 해소하기 위해, 우리는 사회 과학 분야의 고전적인 이원형 경로 탐색 과제인 '맵 태스크(Map Task)'를 기반으로 구축된, 에이전트 협력을 위한 행동 수준의 정신 모델 주석 데이터셋인 ALMANAC을 제시합니다. ALMANAC에는 2,987개의 협업 액션이 포함되어 있으며, 각 액션은 참가자들의 자기 보고 추론, 인식된 파트너 의도 및 인식된 팀 목표를 기록하는 이론 기반의 정신 모델 주석과 함께 제공됩니다. 우리는 ALMANAC을 사용하여 인간의 다음 행동과 정신 모델을 예측하는 데 있어 6개의 LLM을 평가했습니다. 그 결과는 ALMANAC이 모델들이 인간의 협업 행동을 시뮬레이션하고 그 근본적인 정신 모델을 추론하는 능력을 평가하는 데 유용함을 보여줍니다.
Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators. Effective collaboration, however, requires collaborators to continuously maintain and align mental models of their own reasoning,partners' intentions, and shared goals during the collaborative process. Today's agents rarely develop such capabilities since they are primarily optimized for task completion, and the community lacks authentic human collaboration data with action-level mental model annotations that could guide agents toward process-level collaborative competence. To bridge this gap, we present ALMANAC, a dataset of Action-Level Mental model ANnotations for Agent Collaboration built from the Map Task, a classic dyadic routing task from social science. ALMANAC contains 2,987 collaboration actions, each paired with theory-informed mental model annotations that record the participants' self-reasoning, perceived partner intent, and perceived team goal. We benchmark six LLMs on predicting humans' next-turn behavior and mental models. Our results demonstrate ALMANAC's utility in evaluating models' ability to simulate human collaborative behaviors and infer their underlying mental models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.