2607.02141v1 Jul 02, 2026 cs.AI

A$^{2}$utoLPBench: 역 KKT 구성 기반의 자동 생성, 에이전트 친화적 선형 계획법 벤치마크

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction

Yifan Shi
Yifan Shi
Citations: 9,359
h-index: 4
Rongliang Fu
Rongliang Fu
Citations: 124
h-index: 6
Shuo Ren
Shuo Ren
Citations: 164
h-index: 5
Yaohui Han
Yaohui Han
Citations: 9
h-index: 1
Libo Shen
Libo Shen
Citations: 10
h-index: 2
Haodong Lu
Haodong Lu
Citations: 0
h-index: 0
Dongfang Wu
Dongfang Wu
Citations: 0
h-index: 0
Bei Yu
Bei Yu
Citations: 65
h-index: 4
Tsung-Yi Ho
Tsung-Yi Ho
Citations: 122
h-index: 4

대부분의 텍스트 기반 선형 계획법 벤치마크는 사람이 직접 작성하고 레이블링한 정적인 데이터 세트입니다. 이러한 데이터 세트가 공개되면 크기와 난이도가 고정되며, 모든 문제가 향후 LLM(Large Language Model)의 학습 데이터로 유출될 수 있습니다. 본 논문에서는 자연어 텍스트로 작성된 선형 계획법 문제를 사용하여 LLM 기반 에이전트를 테스트하기 위한 벤치마크인 extbf{A$^{2}$utoLPBench}를 제시합니다. 먼저 타당한 점과 이중 변수를 선택하고, 해당 점이 최적점이며 목적 함수 값이 알려진 문제를 구성합니다. 정답은 솔버 호출이나 인간 어노테이터 없이 구성 과정을 통해 얻어지며, 벤치마크 환경에는 참조 솔버-비평기(solver-critic) 기준선과 LLM 기반 에이전트가 읽을 수 있도록 사용 설명이 작성된 Docker 이미지가 포함되어 있습니다. 이러한 환경 덕분에 어떤 에이전트라도 벤치마크를 실행하고 단일 명령어로 보정된 점수를 얻을 수 있습니다. A$^{2}$utoLPBench는 고정된 데이터 세트가 아닌 생성기이기 때문에 다음과 같은 장점을 가지고 있습니다: 무한히 많은 신선한 문제 제공, (n, m) 값으로 조절 가능한 난이도, 구성 과정을 통해 정확하게 도출된 정답, 인간 작성을 대비할 때 문제당 LLM 측 비용이 낮음, 독립적인 배치에서 반복 가능한 점수, 그리고 새로운 시드 범위를 사용할 경우 학습 데이터 유출에 대한 저항성.

Original Abstract

Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is fixed, its difficulty is fixed, and every problem can leak into the training data of future LLMs. We present \textbf{A$^{2}$utoLPBench}, a benchmark for testing LLM-driven agents on linear programming problems written in plain text. We first pick a feasible point and dual, then write down a problem for which that point is optimal and the objective value is known. The answer is known by construction, with no solver call and no human annotator. The evaluation environment bundles a reference solver-critic baseline and a Docker image whose usage instructions are written for an LLM-driven agent to read. With these in place, any agent can run the benchmark and get a calibrated score with one command. Because the benchmark is a generator rather than a fixed dataset, it has properties no fixed dataset can match: an unlimited supply of fresh problems, a difficulty knob set by $(n,m)$, ground-truth answers correct by construction, low LLM-side cost per problem relative to human authoring, repeatable scores across independent batches, and resistance to training-data leakage when fresh post-cutoff seed ranges are used.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!