Solver-in-the-Loop: 자기 수정 및 운영 연구에서의 합리적인 행동을 위한 MDP 기반 벤치마크
Solver-in-the-Loop: MDP-Based Benchmarks for Self-Correction and Behavioral Rationality in Operations Research
운영 연구 실무자들은 비타당성 모델을 반복적인 과정을 통해 디버깅합니다. 이는 불가능한 부분 시스템(IIS)을 분석하고, 제약 조건 충돌을 식별하며, 타당성을 달성할 때까지 체계적으로 모델을 수정하는 과정을 포함합니다. 그러나 기존의 LLM 벤치마크는 운영 연구를 일회성 번역으로 평가합니다. 즉, 문제 설명을 기반으로 솔버 코드를 생성하는 것입니다. 이는 진단 루프를 완전히 무시하는 것입니다. 우리는 평가 루프에 솔버를 포함시키는 두 가지 벤치마크를 소개합니다. extbf{ exttt{ORDebug}}은 9가지 오류 유형을 포괄하는 5,000개 이상의 문제에 대한 반복적인 자기 수정 기능을 평가합니다. 각 수정 작업은 솔버 재실행 및 IIS 재계산을 트리거하여 결정적이고 검증 가능한 피드백을 제공합니다. extbf{ exttt{ORBias}}은 2,000개의 새로운 판매자 인스턴스(1,000개 ID + 1,000개 OOD)를 통해 행동 합리성을 평가하고, 폐쇄형 최적 정책으로부터의 체계적인 편차를 측정합니다. 26개의 모델과 12,000개 이상의 샘플을 분석한 결과, 도메인별 강화 학습 기반 변환(RLVR) 훈련을 통해 80억 개의 매개변수를 가진 모델이 최첨단 API를 능가하는 것으로 나타났습니다. 회복률은 95.3% (86.2% 대비 +9.1%), 진단 정확도는 62.4% (47.8% 대비 +14.6%), 해결에 필요한 단계 수는 2.25회 (3.78회 대비, 1.7배 빠름)로 나타났습니다. extbf{ exttt{ORBias}}에서 교과 과정 기반 훈련은 평가된 모델 중에서 유일하게 ID에서 OOD로의 부정적인 편향 변화(-9.6%)를 달성하여, 체계적인 편향을 48%(20.0%에서 10.4%로 감소)시켰습니다. 이러한 결과는 검증 가능한 오라클을 사용한 프로세스 수준 평가가 규모보다 더 효과적인 표적 훈련을 가능하게 한다는 것을 보여줍니다.
Operations Research practitioners routinely debug infeasible models through an iterative process: analyzing Irreducible Infeasible Subsystems (\IIS{}), identifying constraint conflicts, and systematically repairing formulations until feasibility is achieved. Yet existing LLM benchmarks evaluate OR as one-shot translation -- given a problem description, generate solver code -- ignoring this diagnostic loop entirely. We introduce two benchmarks that place the \textbf{solver in the evaluation loop}. \textbf{\ORDebug{}} evaluates iterative self-correction through 5,000+ problems spanning 9 error types; each repair action triggers solver re-execution and \IIS{} recomputation, providing deterministic, verifiable feedback. \textbf{\ORBias{}} evaluates behavioral rationality through 2,000 newsvendor instances (1,000 ID + 1,000 OOD), measuring systematic deviations from closed-form optimal policies. Across 26 models and 12,000+ samples, we find that domain-specific RLVR training enables an 8B model to surpass frontier APIs: 95.3\% vs 86.2\% recovery rate (+9.1\%), 62.4\% vs 47.8\% diagnostic accuracy (+14.6\%), and 2.25 vs 3.78 steps to resolution (1.7$\times$ faster). On \ORBias{}, curriculum training achieves the only negative ID$\rightarrow$OOD bias drift among models evaluated (-9.6\%), reducing systematic bias by 48\% (from 20.0\% to 10.4\%). These results demonstrate that process-level evaluation with verifiable oracles enables targeted training that outperforms scale.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.