2605.28158v1 May 27, 2026 cs.AI

OR-Space: 산업 최적화 에이전트를 위한 전체 생명주기 워크스페이스 벤치마크

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Chenyue Zhou
Chenyue Zhou
Citations: 6
h-index: 2
Jianghao Lin
Jianghao Lin
Citations: 43
h-index: 2
Jiangyue Zhao
Jiangyue Zhao
Citations: 230
h-index: 9
Dongdong Ge
Dongdong Ge
Citations: 77
h-index: 4
Yinyu Ye
Yinyu Ye
Citations: 225
h-index: 9

대규모 언어 모델(LLM) 기반 에이전트들이 운영 연구(OR) 모델링을 지원하는 데 점점 더 많이 사용되고 있지만, 기존의 OR 관련 벤치마크는 종종 평가를 자체적으로 완결된 문제 설명을 수학적 공식 또는 솔버 프로그램으로 번역하는 단일 단계 작업으로 제한합니다. 이러한 설정은 실제 산업용 OR 워크플로우의 두 가지 특징인 지속적인 다중 아티팩트 워크스페이스와 다단계 작업 생명주기를 간과합니다. 우리는 모델 구축, 모델 수정 및 근거 기반 설명을 평가하기 위한 전체 생명주기 워크스페이스 벤치마크인 OR-Space를 소개합니다. 각 인스턴스는 비즈니스 문서, 구조화된 데이터, 선택적 코드 아티팩트, 솔버 출력 및 작업별 평가기를 포함하는 실행 가능한 워크스페이스입니다. OR-Space는 세 가지 작업 모드를 정의합니다. 'Build'는 에이전트가 이기종 아티팩트로부터 솔버 사용 가능한 최적화 모델을 구축하는 방식, 'Revise'는 에이전트가 변화하는 요구 사항 또는 솔버 피드백에 따라 기존 모델을 수정하면서 유효한 이전 로직을 유지하는 방식, 그리고 'Explain'은 에이전트가 워크스페이스 아티팩트에 흩어져 있는 증거를 사용하여 솔루션, 제약 조건 및 비즈니스 의미에 대한 근거 기반 질문에 답하는 방식을 평가합니다. OR-Space는 지속적인 워크스페이스와 생명주기 지향 작업을 결합하여 에이전트가 엔드투엔드 텍스트 생성 이상의 신뢰할 수 있는 최적화 작업을 수행할 수 있는지 평가합니다. 우리는 벤치마크 설계, 평가 프로토콜 및 품질 관리 파이프라인을 설명하고, OR-Space를 산업용 OR 워크플로우에서 LLM 에이전트의 신뢰성, 실패 모드 및 실제 적용 가능성을 연구하기 위한 벤치마크로 제시합니다.

Original Abstract

Large language model (LLM) agents are increasingly used to assist with operations research (OR) modeling, yet existing OR-oriented benchmarks often reduce evaluation to one-shot translation from a self-contained problem statement into a mathematical formulation or solver program. Such settings abstract away two characteristics of real industrial OR workflows: persistent multi-artifact workspaces and multi-stage task lifecycles. We introduce OR-Space, a full-lifecycle workspace benchmark for evaluating industrial optimization agents across model construction, model revision, and grounded explanation. Each instance is an executable workspace containing business documents, structured data, optional code artifacts, solver outputs, and task-specific evaluators distributed across interdependent files. OR-Space defines three task modes: Build, where agents construct solver-ready optimization models from heterogeneous artifacts; Revise, where agents modify existing models under changing requirements or solver feedback while preserving valid prior logic; and Explain, where agents answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts. By combining persistent workspaces with lifecycle-oriented tasks, OR-Space evaluates whether agents can perform reliable optimization work beyond end-to-end text generation. We describe the benchmark design, evaluation protocol, and quality-control pipeline, and position OR-Space as a benchmark for studying the reliability, failure modes, and practical readiness of LLM agents in industrial OR workflows.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!