Max@K 정책 경사법에 대한 이점 추정 연구
On Advantage Estimates for Max@K Policy Gradients
검증 가능한 보상을 사용하는 강화 학습은 사후 학습 추론 모델에 널리 사용되지만, 희소한 결과 보상은 탐색을 어렵게 만듭니다. 보완적인 접근 방식은 pass@K 및 max@K과 같은 추론 시간 목표를 직접 최적화하는 것입니다. 그러나 이러한 목표를 위한 기존 정책 경사법 추정기는 서로 다른 신호, 기준선 및 정규화를 사용하므로 이들 간의 관계가 불분명합니다. 우리는 기준선 설계 및 이점 중심화를 통해 이 문제를 연구합니다. 해당 분야의 선도적인 방법의 이점 추정기에서 시작하여, 이는 정책 경사법에 편향되지 않지만 중심화되지 않은 이점을 산출한다는 것을 보여줍니다. 그런 다음, 정책 경사법의 편향성을 유지하면서 실현된 배치 이점이 정확하게 중심화되도록 하는 Leave-Two-Out 기준선을 도입합니다. 결과적으로 나온 방법인 MaxPO는 효율적인 이차 시간 구현을 가지며 LLM 사후 학습을 위한 그룹 기반 강화 학습에 자연스럽게 통합됩니다. 또한 max@K의 표준적인 유한 배치 이점을 도출하여 기존 이점 추정기에 대한 통일된 시각을 제공합니다. 실험적으로, L2O 기준선이 기울기 분산을 줄이고 중심화되지 않은 대안보다 성능이 우수하다는 것을 확인했습니다.
Reinforcement learning with verifiable rewards is widely used for post-training reasoning models, but sparse outcome rewards make exploration difficult. A complementary approach is to optimize inference-time objectives such as pass@K and max@K directly, yet existing policy-gradient estimators for these objectives use different signals, baselines, and normalizations, making their relationships unclear. We study this issue through baseline design and advantage centering. Starting from the advantage estimator of a leading method in the field, we show that it is policy-gradient unbiased but yields a non-centered advantage. We then introduce a Leave-Two-Out baseline that preserves policy-gradient unbiasedness while making realized batch advantages exactly centered. The resulting method, MaxPO, has an efficient quadratic-time implementation and integrates naturally into group-based RL for LLM post-training. We further derive the canonical finite-batch advantage for max@K, providing a unified view of existing advantage estimators. Empirically, we verify that the L2O baseline reduces gradient variance and outperforms non-centered alternatives.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.