다중 에이전트 시스템의 장점에 대한 착각
The Illusion of Multi-Agent Advantage
현재 통념은 다중 에이전트 시스템(MAS)이 단일 에이전트 시스템(SAS)보다 우수하다고 주장하며, 컨텍스트 보호, 병렬 처리 및 분산 의사 결정과 같은 이점을 강조합니다. 그러나 이러한 주장에 대한 경험적 증거는 주로 SAS 기준선과의 비교를 통해 이루어지며, 이러한 비교는 고립된 추론 작업을 우선시하는 벤치마크를 사용하므로 MAS의 실제적인 장점을 제대로 평가하지 못합니다. 본 연구에서는 수동으로 설계된 시스템보다 일반화 성능이 향상되도록 설계된 자동 생성 MAS에 초점을 맞춰 SAS, 특히 Chain-of-Thought with Self-Consistency (CoT-SC)와의 엄격하고 체계적인 비교 분석을 수행했습니다. 전통적인 추론 데이터 세트 및 BrowseComp-Plus와 같은 상호 작용형 다단계 워크플로우 작업에서, 자동 생성 MAS는 CoT-SC보다 지속적으로 성능이 낮으며, 비용은 최대 10배 더 높다는 것을 확인했습니다. 이러한 실패 원인을 작업 구조 자체의 한계에서 비롯된 것인지 분리하기 위해, 명시적인 작업 분해, 컨텍스트 분리 및 병렬화 가능성을 특징으로 하는 MAS에 적합한 진단용 합성 데이터 세트를 도입했습니다. 이 데이터 세트에서는 전문 설계자가 설계한 MAS가 자동 생성된 아키텍처보다 성능과 비용 효율성 모두에서 일관되게 우수한 것으로 나타났습니다. 이는 기존 평가 프레임워크가 증가된 계산 비용으로 인한 미미한 효용 가치를 고려하지 못함으로써 복잡한 MAS의 중요한 아키텍처 격차 및 비효율성을 감추고 있음을 보여줍니다. 더욱이, 생성된 MAS 아키텍처에 대한 체계적인 분석 결과, 현재 자동 설계 패러다임은 기능적 유용성과 연결되지 않는 피상적인 복잡성을 우선시하는 아키텍처를 만들어내며, 이는 다중 에이전트 원칙과의 근본적인 불일치를 드러냅니다.
Prevailing wisdom posits that Multi-Agent Systems (MAS) are superior to Single-Agent Systems (SAS), citing advantages like context protection, parallel processing and distributed decision-making. However, empirical support for this claim relies primarily on comparisons with SAS baselines using benchmarks that prioritize isolated reasoning tasks, which do not adequately assess these advantages. Focusing on automatically generated MAS that are designed for enhanced generalizability over manually-designed counterparts, we perform a rigorous, systematic evaluation against SAS, specifically Chain-of-Thought with Self-Consistency (CoT-SC). Across traditional reasoning datasets and tasks with interactive multi-step workflows (e.g., BrowseComp-Plus), we demonstrate that automatic MAS consistently underperform CoT-SC despite being up to 10x more expensive. To isolate these failures from limitations inherent to task structure, we introduce a diagnostic synthetic dataset tailored for MAS featuring explicit task decomposition, context separation and parallelization potential. We show that expert-architected MAS consistently outperforms automatically generated architectures in both raw performance and cost-efficiency on this dataset, demonstrating that existing evaluation frameworks mask critical architectural gaps and inefficiencies of complex MAS by failing to account for the marginal utility of increased computational cost. Critically, systematic deconstruction of the generated MAS architectures reveals that current automated design paradigms produce architectural bloat that prioritizes superficial complexity which does not translate into functional utility, exposing a fundamental misalignment with multi-agent principles.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.