CoffeeBench: 다양한 다중 에이전트 경제 시스템에서 장기적인 LLM 에이전트 성능 평가
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
LLM (Large Language Model) 에이전트가 점차 더 복잡하고 장기간에 걸친 작업을 수행할 수 있게 되면서, 이러한 에이전트의 성능을 경제 시스템 내에서 평가하는 것이 점점 중요해지고 있습니다. 기존 벤치마크는 주로 단일 에이전트가 수동적인 환경과 상호 작용하는 것을 평가하지만, 경제 시스템은 본질적으로 다중 에이전트 시스템으로, 자율적인 에이전트들이 자신의 목표를 추구하면서 통신하고, 협상하며, 장기간에 걸쳐 거래해야 합니다. 우리는 CoffeeBench라는 벤치마크를 소개합니다. CoffeeBench는 다양한 기업으로 구성된 장기적인 다중 에이전트 경제 시스템에서 LLM 에이전트를 평가하기 위한 도구입니다. CoffeeBench에서는 두 명의 농부, 두 명의 로스터, 그리고 두 명의 소매업자가 90일 동안 자율적으로 사업을 운영하며, 통신과 거래를 통해 누적 순수익을 극대화하고 현금, 재고 및 가격 관리를 수행합니다. 평가 대상 모델은 커피 로스터 한 곳을 제어하며, 나머지 기업은 고정된 기준 에이전트에 의해 제어됩니다. 최근 공개된 LLM과 독점적인 LLM을 포함한 여러 모델에서, 모든 모델이 아무런 조치도 취하지 않는 수동적 기준 모델보다 더 나은 성능을 보였으며, 대부분의 모델이 긍정적인 순수익을 달성했습니다. 에이전트 행동 분석 결과, 장기적인 경제 상호 작용에 상당한 차이가 있음을 알 수 있었습니다. 더 높은 성능을 보이는 모델들은 다른 기업들과 더 적극적으로 소통하는 반면, Claude Haiku 4.5는 일시 정지 오류를 보여, 일관성 있는 평가와 계획을 제시함에도 불구하고 반복적으로 아무런 조치를 취하지 않는 경향이 있습니다. 우리는 향후 연구를 지원하기 위해 코드와 에이전트 경로 데이터를 공개합니다.
As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly important. Unlike existing benchmarks that primarily evaluate a single agent interacting with a passive environment, economic systems are inherently multi-agent, requiring autonomous agents to communicate, negotiate, and transact while pursuing their own objectives over extended periods. We introduce CoffeeBench, a benchmark for evaluating LLM agents in a long-horizon multi-agent economy composed of heterogeneous firms. In CoffeeBench, two farmers, two roasters, and two retailers autonomously operate their businesses over a 90-day simulation, each seeking to maximize cumulative net income through communication and transactions while managing cash, inventory, and pricing. The evaluated model controls one coffee roaster, while the remaining firms are controlled by fixed reference agents. Across several recent open-weight and proprietary LLMs, all models outperform a passive baseline that takes no actions, with most achieving positive net income. Analysis of agent behavior reveals substantial differences in long-horizon economic interaction: higher-performing models communicate more actively with other firms, whereas Claude~Haiku~4.5 exhibits an idle-drift failure mode, repeatedly choosing inaction despite producing coherent assessments and plans. We release our code and agent trajectories to support future research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.