2607.18979v1 Jul 21, 2026 cs.AI

무임승차자 제거: 강화 학습을 통한 병렬 추론의 Shapley 기반 보상 할당

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

Xinke Jiang
Xinke Jiang
Citations: 109
h-index: 7
Haoyu Zhang
Haoyu Zhang
Citations: 4
h-index: 1
Yuhan Pan
Yuhan Pan
Citations: 0
h-index: 0
Zhipeng Qiao
Zhipeng Qiao
Citations: 0
h-index: 0
Tao Feng
Tao Feng
Citations: 0
h-index: 0
Wentao Zhang
Wentao Zhang
Citations: 93
h-index: 5
Yuxuan Cheng
Yuxuan Cheng
Citations: 0
h-index: 0
Miao Li
Miao Li
Citations: 50
h-index: 3
Zhen Tao
Zhen Tao
Citations: 29
h-index: 2
Dengji Zhao
Dengji Zhao
Citations: 557
h-index: 14

대규모 언어 모델(LLM)은 다단계 추론에 뛰어나지만, 현재의 병렬 추론 방식은 종종 개별 추론 경로의 기여도를 구별하는 데 실패합니다. 많은 경로가 중복되거나 오해를 불러일으키거나 심지어 해로울 수 있지만, 결과 수준의 보상은 동일한 보상을 할당하여 모호한 학습 신호를 유발하고 불안정한 훈련을 초래합니다. 본 논문에서는 다중 경로 추론에서 세밀한, 경로 수준의 기여도를 부여하는 강화 학습 프레임워크인 Parallel Shapley를 제안합니다. 각 경로를 협력 게임의 플레이어로 간주하고, Shapley 값을 활용하여 한계 기여도를 정량화하며, 생성적 보상 모델을 사용하여 경로 유틸리티를 평가하고, 효율적인 근사화를 위해 몬테카를로 샘플링을 사용합니다. 수학적 추론 벤치마크에 대한 실험 결과, Parallel Shapley는 기존의 기본 모델보다 우수한 성능을 보이며 더 안정적이고 해석 가능한 훈련을 제공하는 것으로 나타났습니다. 본 프레임워크는 효과적으로 "무임승차자"를 제거하고, 보상을 비례적으로 할당하여 LLM에서 다중 경로 추론 능력을 향상시킵니다.

Original Abstract

Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to ambiguous learning signals and unstable training. We propose Parallel Shapley, a reinforcement learning framework that attributes fine-grained, path-level contributions in multi-path reasoning. Treating each path as a player in a cooperative game, we leverage Shapley values to quantify marginal contributions, using a generative reward model to evaluate path utilities and Monte Carlo sampling for efficient approximation. Experiments on mathematical reasoning benchmarks show that Parallel Shapley outperforms existing baselines while providing more stable and interpretable training. Our framework effectively "fishes out the free riders," assigning reward proportionally and improving multi-path reasoning in LLMs.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!