2605.26958v1 May 26, 2026 cs.CL

Tournament-GRPO: 그룹 기반 토너먼트 보상을 활용한 개방형 장문 생성에서의 강화 학습

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation

Yan Gao
Yan Gao
Citations: 44
h-index: 4
Yiqun Chen
Yiqun Chen
Citations: 219
h-index: 8
Erhan Zhang
Erhan Zhang
Citations: 91
h-index: 4
Jiaxin Mao
Jiaxin Mao
Citations: 159
h-index: 7
Xiaochi Wei
Xiaochi Wei
Citations: 11
h-index: 2
Yi Wu
Yi Wu
Citations: 28
h-index: 4
Yao Hu
Yao Hu
Citations: 16
h-index: 2
Wei Yang
Wei Yang
Citations: 6
h-index: 1
Zixuan Yang
Zixuan Yang
Citations: 18
h-index: 3
Zihang Shen
Zihang Shen
Citations: 0
h-index: 0

개방형 장문 생성 환경에서의 강화 학습은 신뢰할 수 있는 정답 데이터와 자동 평가 지표가 부족하기 때문에 어려운 과제입니다. 기존의 rubic 기반 방법들은 주로 pointwise 방식으로 LLM을 활용하여 점수를 매기지만, 이러한 절대적인 점수는 복잡한 응답에 대해 정확하게 조정하기 어렵고, 동일한 질문에 대한 여러 결과물 간의 구별력을 약화시킬 수 있으며, 최적화 과정에서 포화될 가능성이 있습니다. 본 연구에서는 Tournament-GRPO라는 그룹 기반 보상 프레임워크를 제안합니다. 이 방법은 rubic 가이드라인에 따라 LLM이 제공하는 판단을 반복적인 다단계 토너먼트를 통해 상대적인 보상으로 변환합니다. Tournament-GRPO는 동일한 질문에 대한 여러 후보들을 그룹별로 비교하고, 토너먼트 결과를 누적하여 GRPO 학습을 위한 그룹별 보상을 생성합니다. Deep Research Bench에서의 실험 결과, Tournament-GRPO는 기존의 보상 설계 방법들보다 일관되게 우수한 성능을 보이며, 가장 강력한 기준 모델 대비 4.52점의 전반적인 점수 향상을 달성했습니다. 추가 분석 결과, 토너먼트 보상은 효과성과 효율성 사이의 유리한 균형을 제공하며, 토너먼트 설계 방식이 학습 동역학에 영향을 미치는 것으로 나타났습니다. 이러한 결과는 rubic 가이드라인 기반의 토너먼트 비교가 개방형 장문 생성 환경에서의 강화 학습을 위한 효과적인 보상 신호를 제공할 수 있음을 시사합니다.

Original Abstract

Reinforcement learning in open-ended long-form generation is challenging because reliable reference answers and automatic metrics are often unavailable. Existing rubric-based methods typically rely on pointwise LLM-as-a-judge scoring, but absolute scores are difficult to calibrate across complex responses, may provide weak discrimination among same-query rollouts, and can become saturated during optimization. We propose Tournament-GRPO, a group-wise reward framework that converts rubric-guided LLM judgments into relative rewards through repeated multi-round tournaments among same-query rollouts. Tournament-GRPO compares candidates within groups, accumulates tournament outcomes, and normalizes them into group-wise rewards for GRPO training. Experiments on Deep Research Bench show that Tournament-GRPO consistently outperforms existing reward-design baselines, achieving a 4.52-point overall-score improvement over the strongest baseline. Further analyses show that tournament rewards provide a favorable effectiveness--efficiency trade-off and that tournament design affects training dynamics. These results suggest that rubric-guided tournament comparison provides an effective reward signal for reinforcement learning in open-ended long-form generation.

2 Citations
0 Influential
3.5 Altmetric
19.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!