PolyWorkBench: 다국어 장기 호라이즌 LLM 에이전트 성능 평가
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
대규모 언어 모델(LLM) 기반 에이전트는 계획, 도구 사용 및 외부 환경과의 상호 작용을 요구하는 장기적인 작업에서 뛰어난 성능을 보여주었습니다. 그러나 대부분의 기존 벤치마크는 암묵적으로 단일 언어 환경을 가정하며, 추론, 도구 호출 및 출력 생성과 같은 전체 실행 프로세스가 하나의 언어 내에서 이루어집니다. 반면, 실제 응용 프로그램에서는 종종 통합된 워크플로우 내에서 다국어 입력 및 출력이 사용되지만, 다국어 처리와 에이전트 실행 간의 상호 작용은 아직 충분히 연구되지 않았습니다. 본 논문에서는 다국어 장기 호라이즌 업무 워크플로우에 대한 LLM 에이전트 평가를 위한 벤치마크인 PolyWorkBench를 소개합니다. PolyWorkBench는 상거래, 지식 기반 업무, 법률 분석, 현지화 및 제조를 포함한 다섯 가지 영역에서 총 67개의 작업으로 구성되어 있으며, 에이전트는 이종의 다국어 입력을 처리하고, 반복적인 추론을 수행하며, 외부 도구를 호출하고, 구조화된 출력을 생성해야 합니다. 포괄적인 평가를 위해 우리는 구조적 채점, 실행 가능한 검증 및 LLM 기반 의미 분석을 결합한 하이브리드 프레임워크를 제안합니다. 이러한 설계는 복잡한 워크플로우 전반에 걸쳐 기능적 정확성과 언어적 일관성을 모두 포착할 수 있도록 합니다. 실험 결과에 따르면 최첨단 LLM 에이전트는 다국어 워크플로우 환경에서 단일 언어 환경의 에이전트에 비해 성능 저하가 심각합니다. 우리의 분석은 다국어가 추론 및 실행 단계에서 복합적인 영향을 미치며, 이는 에이전트 평가 시 언어 변동과 절차적 의사 결정을 함께 모델링하는 것의 중요성을 강조한다는 것을 보여줍니다.
Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire execution process, including reasoning, tool invocation, and output generation, is conducted within a single language. In contrast, real-world applications often involve multilingual inputs and outputs within a unified workflow, yet the interaction between multilinguality and agentic execution remains underexplored. In this work, we introduce PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows. PolyWorkBench consists of 67 tasks across five domains, including commerce, knowledge work, legal analysis, localization, and manufacturing, where agents must process heterogeneous multilingual inputs, perform iterative reasoning, invoke external tools, and produce structured outputs. To enable comprehensive evaluation, we propose a hybrid framework that combines structural grading, executable verification, and LLM-based semantic assessment. This design allows us to capture both functional correctness and linguistic consistency across complex workflows. Empirical results show that state-of-the-art LLM agents suffer significant performance degradation in multilingual workflow settings compared to monolingual counterparts. Our analysis suggests that multilinguality introduces compounding effects across reasoning and execution steps, highlighting the importance of jointly modeling language variation and procedural decision-making in agent evaluation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.