2608.02149v1 Aug 03, 2026 cs.AI

평균을 넘어: LLM 추론을 위한 다중 모멘트 기반 정책 최적화

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning

Luoyi Fu
Luoyi Fu
Citations: 3,774
h-index: 32
Jiaxin Ding
Jiaxin Ding
Citations: 166
h-index: 8
Fan Xu
Fan Xu
Citations: 10
h-index: 2
Yule Xie
Yule Xie
Citations: 25
h-index: 2
Xin Ding
Xin Ding
Citations: 52
h-index: 3
Yijun Zhang
Yijun Zhang
Citations: 0
h-index: 0
Haoxiang Zhang
Haoxiang Zhang
Citations: 72
h-index: 1

강화 학습은 대규모 언어 모델(LLM)의 추론 능력을 향상시키는 핵심적인 패러다임으로 자리 잡았습니다. 기존 방법들은 일반적으로 다양한 문제에서 발생하는 실패 확률을 줄이는 것을 목표로 합니다. 본 논문에서는 LLM 추론을 위한 정책 최적화에 대한 모멘트 기반 관점을 제시합니다. 임의로 선택된 문제의 실패 확률을 확률 변수로 간주하고, 이 확률 변수의 모멘트를 통해 최적화 목표를 정의합니다. 이러한 관점에서 많은 기존 방법들이 실패 확률 분포의 단일 모멘트만을 최적화하며, 분포의 전체적인 구조는 대부분 고려되지 않습니다. 우리는 다중 모멘트 기반 정책 최적화(MMPO)라는 새로운 정책 최적화 프레임워크를 제안합니다. MMPO는 실패 확률 분포의 여러 개의 모멘트를 동시에 최소화합니다. MMPO는 첫 번째 성공적인 응답을 얻기까지 예상되는 절단된 시간을 최소화하는 직접적인 해석을 가집니다. MMPO 외에도, 우리는 다양한 모멘트 프로필을 체계적으로 유도하고 더 넓은 범위의 정책 최적화 목표에 대한 통합적인 관점을 제공하는 일반적인 모멘트 변환 프레임워크를 개발했습니다. 다섯 가지 수학적 추론 벤치마크와 다양한 크기의 모델에 대한 실험 결과, MMPO는 강력한 기준 방법을 지속적으로 능가한다는 것을 보여줍니다. 우리는 본 논문에서 제시하는 모멘트 기반 관점이 LLM 추론을 위한 정책 최적화 목표 설계에 새로운 통찰력을 제공할 수 있기를 바랍니다.

Original Abstract

Reinforcement learning has become a central paradigm for improving the reasoning capabilities of large language models. Existing methods generally aim to reduce the failure probabilities induced across problems. In this paper, we introduce a moment-based perspective on policy optimization for LLM reasoning by treating the failure probability of a randomly sampled problem as a random variable and characterizing optimization objectives through its moments. Under this perspective, many existing methods optimize only a single moment of the failure-probability distribution, leaving its broader distributional structure largely uncharacterized. We propose \textbf{M}ulti-\textbf{M}oment \textbf{P}olicy \textbf{O}ptimization (MMPO), a novel policy optimization framework that jointly minimizes multiple moments of the failure-probability distribution. MMPO admits a direct operational interpretation as minimizing the expected truncated time required to obtain the first successful response. Beyond MMPO, we further develop a general moment-transformation framework that systematically induces different moment profiles and provides a unified view of a broader family of policy optimization objectives. Experiments across five mathematical reasoning benchmarks and models of different scales demonstrate that MMPO consistently outperforms strong baselines. We hope this moment-based perspective offers new insights into the design of policy optimization objectives for LLM reasoning.

0 Citations
0 Influential
16 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!