다국어, 다중 에이전트 계획 실패에 대한 실용적인 진단
An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures
다국어 다중 에이전트 시스템은 영어 환경을 넘어 상당한 성능 저하를 보이지만, 기존 연구에서는 사용자 요청이 실행 가능한 계획으로 변환되는 과정에서 어떤 중요한 정보가 손실되는지 명확하게 밝히는 경우가 드물었습니다. 본 연구에서는 다중 에이전트 시스템의 플래너를 요청-액션 인터페이스로 보고, 실제 작업 수행 실패 사례를 분석하여 계획 수립 오류에 대한 실용적인 분류 체계를 도출했습니다. LLM 기반 분석 결과, 이러한 오류는 언어 자원 가용성이 감소함에 따라 성공하지 못하는 실행 비율을 증가시키며, 특히 저자원 언어에서 더욱 두드러진 영향을 미치는 것으로 나타났습니다. 제안된 분류 체계가 개선에 도움이 되는지 검증하기 위해, 플래너와 하위 에이전트에 분류 체계의 핵심 요소를 명시적으로 전달하는 TART (Taxonomy-Guided Actionable Representation)를 도입했습니다. 다양한 언어, 세 가지 LLM 모델 아키텍처, 두 개의 데이터셋, 그리고 두 가지 에이전트 구성 환경에서 TART는 일관되게 성능 향상을 보였습니다. 특히 다국어 GAIA 데이터셋에서, TART는 저자원부터 고자원 환경에 이르는 열한 개 언어 전반에 걸쳐 최고 수준의 시스템 정확도를 5.6%p만큼 향상시켰습니다.
Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is lost when user requests are converted into executable plans. We study the planner in a multi-agent system as the request-to-action interface and derive an actionable taxonomy of planning-grounding failures from failed real-world task executions. LLM-based analysis shows that these failures constitute an increasing share of unsuccessful executions as language-resource availability declines, with the strongest effects in low-resource languages. To test whether the taxonomy supports mitigation, we introduce TART, Taxonomy-Guided Actionable Representation, that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents. Across multiple languages, three LLM backbones, two datasets, and two agentic configurations, TART consistently improves performance. On multilingual GAIA, it raises a state-of-the-art system's accuracy by 5.6 percentage points averaged across eleven languages spanning low- to high-resource settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.