리더보드 너머: 거대 언어 모델 에이전트의 도구 사용, 계획, 그리고 추론 실패에 대한 종합 분석
Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents
거대 언어 모델(LLM) 에이전트는 도구 사용 능력, 다단계 작업 계획 수립 능력, 다른 에이전트와의 협업 능력, 그리고 장기적인 관점에서 작동하는 능력을 평가받고 있습니다. 보고되는 성능 향상은 종종 다양한 평가 시도에서 반복적으로 나타나는 실패 사례들을 가립니다. 본 논문은 2023년부터 2026년 사이에 발표된 27개의 벤치마크, 분류 체계, 그리고 감사 관련 연구(총 19개의 개별 벤치마크)를 종합하여 에이전트의 한계를 포괄적으로 분석한 분류 체계를 제시합니다. 우리가 알고 있는 바로는, 본 논문은 도구 사용, 계획 수립, 장기 추론, 다중 에이전트 협업, 안전성, 그리고 측정의 타당성을 통합하여 LLM 에이전트의 한계에 대한 단일하고 통일적인 분류 체계를 처음으로 제시하는 연구입니다. 우리는 여섯 가지 주요 실패 유형을 다음과 같이 정의했습니다: (1) 도구 호출 및 매개변수 수준 오류, (2) 계획 수립 및 제약 조건 만족 실패, (3) 컨텍스트 누적에 따른 장기 추론 능력 저하, (4) 다중 에이전트 협업 실패, (5) 적대적인 또는 불명확한 상황에서의 안전 및 보안 실패, 그리고 (6) 측정의 타당성 문제. 본 분류 체계는 독립적으로 보고된 오류 범주들을 에이전트의 사고-행동 과정의 다양한 단계에 해당하는 주제로 그룹화하여 반복적으로 도출되었습니다. 문헌 분석 결과, 작업 길이가 증가함에 따라 오류가 비선형적으로 누적되며, 개별 하위 작업에서의 강력한 성능이 전체적인 성공으로 이어지지 않는 경우가 많고, 추가적인 지원 요소가 항상 안정성을 향상시키지는 못하는 것으로 나타났습니다. 반면, 단일 턴 도구 사용, 짧은 기간 웹 탐색, 그리고 제한된 범위의 코딩 작업에서는 상당한 진전이 있었습니다.
Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring failure modes documented across otherwise unrelated evaluation efforts. This paper synthesizes 27 benchmark, taxonomy, and audit papers (2023-2026), spanning 19 distinct benchmarks, into a cross-cutting taxonomy of agent limitations. To our knowledge, this is the first synthesis that integrates evidence across tool use, planning, long-horizon reasoning, multi-agent coordination, safety, and measurement validity into a single, unified taxonomy of LLM agent limitations. We identify six failure clusters: (1) tool invocation and parameter-level errors, (2) planning and constraint-satisfaction failures, (3) long-horizon degradation from context accumulation, (4) multi-agent coordination failures, (5) safety and security failures under adversarial or underspecified conditions, and (6) measurement validity problems. The taxonomy was derived iteratively by grouping independently reported error categories into themes corresponding to distinct stages of the agent reasoning-to-action pipeline. Across the literature, we find that failures compound nonlinearly with task length, that strong performance on individual sub-tasks does not reliably translate into end-to-end success, and that additional scaffolding does not consistently improve reliability. At the same time, substantial progress has been demonstrated in single-turn tool use, short-horizon web navigation, and narrowly scoped coding tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.