2606.31174v1 Jun 30, 2026 cs.AI

ClawArena-Team: 언어 모델 에이전트에서의 서브에이전트 오케스트레이션 및 동적 워크플로우 벤치마킹

ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

Kai Xiong
Kai Xiong
Citations: 93
h-index: 5
Cihang Xie
Cihang Xie
Citations: 156
h-index: 6
Huaxiu Yao
Huaxiu Yao
Citations: 218
h-index: 8
Shi Qiu
Shi Qiu
Citations: 93
h-index: 6
Haonian Ji
Haonian Ji
Citations: 80
h-index: 5
Xinyue Ye
Xinyue Ye
Citations: 41
h-index: 2
Zeyu Zheng
Zeyu Zheng
Citations: 5,358
h-index: 20

최근 대규모 언어 모델(LLM) 에이전트는 독립적인 문제 해결 도구로 사용되는 대신, 관리자 역할을 수행하는 경향이 있습니다. 즉, 주 모델이 특화된 서브에이전트를 생성하고 작업 위임을 하며, 동적 워크플로우를 통해 이들의 병렬 및 비동기 응답을 조정합니다. 하나의 모델이 실제로 이러한 팀을 운영할 수 있는지 여부는 아직 제대로 측정되지 않았습니다. 기존 벤치마크는 특정 정책의 문제 해결 능력 또는 고정된 다중 에이전트 시스템의 창발적 행동을 평가하지만, 리더 역할을 하는 단일 LLM의 관리 능력을 별도로 평가하는 것은 아닙니다. 본 연구에서는 이러한 관리 능력을 측정하기 위한 벤치마크인 ClawArena-Team을 소개합니다. ClawArena-Team은 258개의 평가 라운드와 72단계 업데이트를 포함하는 41개의 다중 회차, 멀티모달, 멀티 디렉토리 시나리오로 구성되어 있습니다. 주 에이전트는 의도적으로 제약 조건을 갖도록 설계되었습니다. 즉, 텍스트만 인식하고 작업 공간의 일부에만 직접 액세스할 수 있습니다. 또한, 고정된 로컬 서브에이전트 풀을 사용하므로, 점수 차이는 단순히 모델의 성능이 아닌 관리 기술을 반영합니다. 모든 평가는 LLM 평가자 없이 실행 기반으로 이루어지며, 전체 점수인 '서브에이전트 관리 점수(SMS)'는 작업 정확도에 최소 권한 및 모달리티 라우팅 요소를 곱하여 계산됩니다. 12개의 독점 모델, 커뮤니티 호스팅 모델, 자체 호스팅 모델을 대상으로 실험한 결과, 관리 병목 현상은 인식 능력보다는 권한 부여 문제임을 확인했습니다 (어떤 모델도 작업 공간에 대한 권한 정확도가 50%를 초과하지 않음). 또한, 비용과 관리 품질은 서로 연관되지 않는다는 것을 보여주었습니다 (API 비용은 최대 100배까지 차이가 있지만, 전체 점수는 최대 4배 정도만 차이) 또한 가장 저렴한 오픈 소스 모델들이 파레토 최적 지점에 위치하며, 대부분의 리더보드 점수가 9.9점 이내로 밀집되어 있는 반면, 오케스트레이션 행동은 최대 10배 이상 차이가 나는 것을 확인했습니다. 코드와 데이터는 공개될 예정입니다.

Original Abstract

Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows. Whether one model can actually run such a team is largely unmeasured: existing benchmarks score a policy's own task-solving or a fixed multi-agent system's emergent behavior, but none isolate the management ability of the single LLM acting as leader. We introduce ClawArena-Team, a benchmark of 41 multi-turn, multimodal, multi-directory scenarios spanning 258 evaluation rounds and 72 staged updates that measures this management ability. The main agent is deliberately constrained: it natively perceives only text and directly accesses only part of the workspace. It commands a fixed, locally served subagent pool, so score differences reflect management skill, not raw capability. All scoring is execution-based with no LLM judge: an overall score -- the Subagent-Management Score (SMS) -- multiplies task correctness by a least-privilege and modality-routing factor. Across twelve proprietary, community-hosted, and self-hosted models, experiments show that the management bottleneck is privilege granting rather than perception (no model exceeds 50% workspace-permission precision); that cost and management quality are decoupled (API cost spans over 100 times while the overall score spans under 4 times, with the cheapest open models on the Pareto frontier); and that most leaderboard scores cluster within a 9.9-point band while orchestration behaviors diverge by more than an order of magnitude. Code and data will be released.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!