무임승차자 제거: 강화 학습을 통한 병렬 추론의 Shapley 기반 보상 할당
Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning
대규모 언어 모델(LLM)은 다단계 추론에 뛰어나지만, 현재의 병렬 추론 방식은 종종 개별 추론 경로의 기여도를 구별하는 데 실패합니다. 많은 경로가 중복되거나 오해를 불러일으키거나 심지어 해로울 수 있지만, 결과 수준의 보상은 동일한 보상을 할당하여 모호한 학습 신호를 유발하고 불안정한 훈련을 초래합니다. 본 논문에서는 다중 경로 추론에서 세밀한, 경로 수준의 기여도를 부여하는 강화 학습 프레임워크인 Parallel Shapley를 제안합니다. 각 경로를 협력 게임의 플레이어로 간주하고, Shapley 값을 활용하여 한계 기여도를 정량화하며, 생성적 보상 모델을 사용하여 경로 유틸리티를 평가하고, 효율적인 근사화를 위해 몬테카를로 샘플링을 사용합니다. 수학적 추론 벤치마크에 대한 실험 결과, Parallel Shapley는 기존의 기본 모델보다 우수한 성능을 보이며 더 안정적이고 해석 가능한 훈련을 제공하는 것으로 나타났습니다. 본 프레임워크는 효과적으로 "무임승차자"를 제거하고, 보상을 비례적으로 할당하여 LLM에서 다중 경로 추론 능력을 향상시킵니다.
Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to ambiguous learning signals and unstable training. We propose Parallel Shapley, a reinforcement learning framework that attributes fine-grained, path-level contributions in multi-path reasoning. Treating each path as a player in a cooperative game, we leverage Shapley values to quantify marginal contributions, using a generative reward model to evaluate path utilities and Monte Carlo sampling for efficient approximation. Experiments on mathematical reasoning benchmarks show that Parallel Shapley outperforms existing baselines while providing more stable and interpretable training. Our framework effectively "fishes out the free riders," assigning reward proportionally and improving multi-path reasoning in LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.