ORAgentBench: LLM 에이전트는 복잡한 운영 연구 과제를 완벽하게 해결할 수 있을까요?
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
대규모 언어 모델(LLM)은 실행 가능한 환경에서 다단계 작업을 수행하는 자율적인 에이전트로 점점 더 많이 활용되고 있지만, 이러한 모델들이 실제 운영 연구(OR) 문제를 얼마나 잘 해결할 수 있는지는 아직 명확하지 않습니다. 기존의 OR 평가 방법은 종종 모델링과 문제 해결을 분리하고, 미리 정의된 또는 텍스트 기반의 사례에 의존하며, 운영 관련 자료에서 검증된 결정까지 전체 워크플로우를 테스트하는 경우는 드뭅니다. 본 연구에서는 자율적인 에이전트가 복잡한 엔드투엔드 운영 연구 과제를 수행하는 능력을 평가하기 위한 실행 기반 벤치마크인 ORAgentBench를 소개합니다. ORAgentBench는 다양한 운영 시나리오에 걸쳐 107개의 인간 검토 작업을 포함하며, 각 작업은 자연어 설명, 다중 파일 데이터, 구성 요소 및 필수 제출 스키마가 포함된 격리된 환경으로 패키징되어 있습니다. 에이전트는 솔루션 코드를 작성하고 실행해야 하며, 제출물은 스키마 유효성, 제약 조건 준수 여부 및 정규화된 목표 품질 측면에서 숨겨진 검증기를 통해 평가됩니다. 14개의 최첨단 에이전트 모델 구성에 대한 실험 결과, 현재의 에이전트는 신뢰할 수 있는 OR 실무 수행 능력과 아직 거리가 먼 것으로 나타났습니다. 가장 성능이 좋은 에이전트조차도 전체 작업의 35.51%와 어려운 작업의 20.59%만을 성공적으로 해결했으며, 많은 실행 가능한 제출물은 여전히 요구되는 품질 기준을 충족하지 못했습니다. 추가 분석 결과, 오류는 주로 전략적 약점(예: 누락된 운영 규칙, 취약한 모델링, 미흡한 실행 가능 솔루션 생성 및 불충분한 솔루션 개선)에서 비롯됩니다. OR에 특화된 절차적 기술은 어려운 작업의 성공률을 높이는 데 도움이 되지만, 반드시 솔루션 품질이나 합격률을 향상시키지는 않습니다. 이러한 결과는 OR 에이전트의 발전을 위해서는 단순히 괜찮은 최적화 코드를 작성하는 것을 넘어 신뢰할 수 있고 고품질의 운영 의사 결정을 내리는 방향으로 나아가야 함을 시사합니다.
Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR) work remains unclear. Existing OR evaluations often decouple modeling from solving, rely on pre-formalized or text-only instances, and rarely test the full workflow from operational artifacts to validated decisions. In this work, we introduce ORAgentBench, an execution-grounded benchmark for evaluating autonomous agents on challenging end-to-end operations research tasks. It contains 107 human-reviewed tasks across diverse operational scenarios, each packaged in an isolated environment with a natural-language brief, multi-file data, configuration artifacts, and a required submission schema. Agents must write and run solution code, and their submissions are evaluated by hidden validators for schema validity, hard-constraint feasibility, and normalized objective quality. Experiments with fourteen frontier agent-model configurations show that current agents remain far from reliable OR practice. The best agent passes only 35.51% of all tasks and 20.59% of hard tasks, and many feasible submissions still fall below the required quality threshold. Failure analysis further shows that errors are dominated by strategic weaknesses, including missed operational rules, brittle formulations, weak feasible-solution construction, and insufficient solution improvement. OR-specific procedural skills increase hard-task feasibility, but do not reliably improve solution quality or pass rate. These results suggest that progress in OR agents requires moving beyond plausible optimization code toward dependable, high-quality operational decision-making.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.