2606.25832v1 Jun 24, 2026 cs.LG

MiniOpt: 제한된 자원을 활용하여 일반적인 최적화 문제를 모델링하고 해결하기 위한 추론 기반 접근 방식

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

Bingdong Li
Bingdong Li
Citations: 1,067
h-index: 13
Xiangfeng Wang
Xiangfeng Wang
Citations: 43
h-index: 2
Zixiang Di
Zixiang Di
Citations: 30
h-index: 2
K. Tang
K. Tang
Citations: 139
h-index: 3
Keshi Zhao
Keshi Zhao
Citations: 7
h-index: 2
Hong Qian
Hong Qian
Citations: 1,069
h-index: 12
Xiang Shu
Xiang Shu
Citations: 100
h-index: 2
Yaolin Wen
Yaolin Wen
Citations: 9
h-index: 2
Qitao Shi
Qitao Shi
Citations: 146
h-index: 7
Xingyu Lu
Xingyu Lu
Citations: 3
h-index: 1
Jun Zhou
Jun Zhou
Citations: 95
h-index: 1
Yang Yu
Yang Yu
Citations: 297
h-index: 5

최적화 지향 대규모 언어 모델(LLM)의 경우, 다양한 최적화 문제에 대한 강력한 일반화 성능을 달성하면서도 제한된 학습 자원을 요구하는 것은 여전히 어려운 과제입니다. 기존 방법은 일반적으로 대규모 감독 데이터셋, 비용이 많이 드는 추론 주석, 그리고 값비싼 중간 단계 검증에 의존하여 상당한 학습 오버헤드를 발생시킵니다. 이러한 문제점을 해결하기 위해, 우리는 "추론-모델링-해결" 패러다임을 통해 최적화 문제를 해결하는 강화 학습 프레임워크인 MiniOpt를 제안합니다. MiniOpt는 최적화 추론을 구조화된 최적화 모델링과 실행 가능한 솔버 생성으로 분해합니다. 이러한 패러다임을 기반으로, 우리는 수식 및 해법을 동시에 평가하여 효과적인 정책 학습을 가능하게 하는 계층적 점수 구조를 가진 보상 함수인 OptReward를 도입했습니다. 또한, 탐색 효율성을 향상시키고 소형 모델에 대한 강화 학습의 안정성을 높이는 최적화 지향 정책 최적화 전략을 개발했습니다. 광범위한 실험 결과, MiniOpt-3B는 다양한 최적화 유형, 문제 시나리오 및 작업 영역에서 강력한 최적화 일반화 성능을 보여줍니다. 100억 개 미만의 파라미터를 가진 모델의 경우, MiniOpt 시리즈가 가장 높은 평균 해결 정확도(SA)를 달성했습니다. 100억 개 이상의 파라미터를 가진 모델의 경우에도 MiniOpt는 경쟁력 있는 성능을 유지합니다. 이러한 결과는 최적화 지향적인 보상 설계와 강화 학습이 강력한 최적화 일반화 능력을 갖춘 소형의 최적화 전문 언어 모델을 개발하는 효과적인 방법을 제공한다는 것을 시사합니다. 코드는 https://github.com/Hsiang-1/MiniOpt 에서 이용 가능합니다.

Original Abstract

Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a challenging problem for optimization-oriented large language models (LLMs). Existing approaches typically rely on large-scale supervised datasets, costly reasoning annotations, and expensive intermediate step verification, resulting in substantial training overhead. To address these challenges, we propose MiniOpt, a reinforcement learning framework that learns to solve optimization problems through an "reasoning-to-model-and-solve" paradigm. MiniOpt decomposes optimization reasoning into structured optimization modeling and executable solver generation. Building upon this paradigm, we introduce OptReward, a reward function with hierarchical score structure that jointly evaluates formulation and solution, enabling effective policy learning without expert demonstrations. We further develop an optimization-oriented policy optimization strategy that improves exploration efficiency and stabilizes reinforcement learning for compact models. Extensive experiments show that MiniOpt-3B exhibits strong optimization generalization across various optimization types, problem scenarios, and task domains. For models with fewer than 10B parameters, MiniOpt series achieves the highest average solving accuracy (SA). For models with more than 10B parameters, MiniOpt still shows competitive performance. These results suggest that optimization-oriented reward design and reinforcement learning provide an effective pathway for developing compact optimization-specialized language models with strong optimization generalization capabilities. The code is available at https://github.com/Hsiang-1/MiniOpt.

1 Citations
0 Influential
38.489476363992 Altmetric
6.9 Score
Original PDF
10

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!