2606.17459v1 Jun 16, 2026 cs.AI

LLM이 CEO가 될 수 있을까? 멀티 에이전트 시뮬레이션을 활용한 전략적 자원 재분배 성능 평가

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

Xueqing Peng
Xueqing Peng
Citations: 656
h-index: 13
Lingfei Qian
Lingfei Qian
Citations: 359
h-index: 10
Zhuohan Xie
Zhuohan Xie
Citations: 327
h-index: 9
Yuyang Dai
Yuyang Dai
Citations: 10
h-index: 2

대규모 언어 모델(LLM)의 의사 결정 능력을 평가하는 것은 중요한 연구 분야이지만, 기존 벤치마크는 추론, 지식 검색 및 스타일화된 환경에서의 경제적 합리성과 같은 고립된 인지 과제에 초점을 맞추고 있습니다. 이러한 평가는 실제 경영진 의사 결정의 핵심적인 과제를 간과합니다. 즉, 정보 비대칭, 조직 제약 및 시간 의존성 하에서 전문 이해 관계자로부터 받는 상반되는 권고 사항을 통합하는 것입니다. 본 연구에서는 CEO 수준의 전략적 자원 재분배를 평가하는 멀티 에이전트 벤치마크인 extsc{CEO-Bench}를 소개합니다. extsc{CEO-Bench}는 LLM 에이전트가 다수의 라운드로 구성된 조직 환경 내에서 사업 단위 간에 자본을 재분배하는 과정을 수행하며, 각 에이전트는 사적인 정보와 고유한 우선 순위를 가진 네 명의 역할 기반 C-suite 어드바이저(CFO, CTO, COO, CMO)로부터 상반되는 조언을 받습니다. LLM은 이러한 의견을 종합하여 구체적인 분배 계획을 수립하며, 이 계획은 역할 통합, 상황에 따른 과감함, 역사적 맥락을 고려한 판단 및 계획의 타당성이라는 네 가지 측면에서 평가됩니다. 다섯 개의 최첨단 모델을 대상으로 13가지 시나리오를 통해 실험한 결과, 모든 모델이 높은 수준의 구조적 타당성을 보이지만, 가장 어려운 영역인 전략적 조정 능력에서는 상당한 차이를 나타냅니다. 본 연구는 단일 어드바이저에 대한 과도한 의존, 불확실성 하에서의 보수적인 기본 설정 및 역사적 기억 상실과 같은 체계적인 실패 사례를 파악했으며, 또한 통합성과 과감함 간의 상충 관계가 존재한다는 사실을 밝혀냈습니다. 즉, 상반된 관점에 더 깊이 참여하는 모델일수록 결정력이 떨어지는 경향이 있습니다. 이러한 결과는 조직 의사 결정자로서 LLM의 현재 능력 범위를 명확히 하고, 미래의 AI 기반 경영 지원 시스템 설계에 필요한 정보를 제공합니다.

Original Abstract

Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings. These evaluations overlook the defining challenge of real executive decision-making: integrating conflicting recommendations from specialized stakeholders under information asymmetry, organizational constraints, and temporal dependencies. We introduce \textsc{CEO-Bench}, a multi-agent benchmark that evaluates LLMs on CEO-level strategic resource reallocation -- the process of redirecting capital across business units in a multi-round, constraint-rich organizational environment. In \textsc{CEO-Bench}, LLM agents receive conflicting advice from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO), each with private signals and distinct priorities, and must synthesize these into a concrete allocation plan evaluated along four dimensions: role integration, conditional boldness, history-sensitive judgment, and plan validity. Experiments across five frontier models on 13 scenarios reveal that all models achieve high structural validity but diverge sharply on strategic calibration -- the hardest capability layer. We identify systematic failure modes including single-advisor capture, conservative default under ambiguity, and historical amnesia, and uncover a structural integration-boldness tradeoff: models that engage more deeply with conflicting perspectives tend to produce less decisive action. These findings delineate the current capability boundary of LLMs as organizational decision-makers and inform the design of future AI-assisted executive systems.

1 Citations
0 Influential
6.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!