2605.27846v1 May 27, 2026 cs.AI

EAPO: 정책 최적화를 위한 엔트로피 기반 적응형 양/음 샘플 가중치 부여 방법 - 개방형 질의 응답 시스템에 적용

EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA

Yujing Wang
Yujing Wang
Citations: 23
h-index: 2
Yuwei Miao
Yuwei Miao
Citations: 29
h-index: 3
Gen Li
Gen Li
Citations: 25
h-index: 2
Yunsheng Zeng
Yunsheng Zeng
Citations: 47
h-index: 4
Siyu Chen
Siyu Chen
Citations: 16
h-index: 2
Yu Qiao
Yu Qiao
Citations: 1,704
h-index: 5
Jianwei Lv
Jianwei Lv
Citations: 121
h-index: 4
Bo Yuan
Bo Yuan
Citations: 47
h-index: 2
Xiandong Li
Xiandong Li
Citations: 3
h-index: 1
Luning Wang
Luning Wang
Citations: 362
h-index: 3
Junfeng Wang
Junfeng Wang
Citations: 959
h-index: 6

대규모 추론 모델은 일반적으로 검증 가능한 보상을 이용한 강화 학습(RLVR)을 통해 훈련됩니다. 그러나 기존 방식들은 양성 및 음성 샘플에 대해 고정된 가중치를 사용하며, 이러한 방식이 개방형 질의 응답(QA) 시스템으로 일반화되기 어렵습니다. 본 논문에서는 개방형 QA를 위한 강화 학습에서 양성 및 음성 샘플의 역할을 체계적으로 분석합니다. 우리는 보상 평균 기반 전략을 사용하여 양성 샘플과 음성 샘플을 구별하고, 음성 샘플이 응답 다양성과 성능 상한에 주로 영향을 미치는 반면, 양성 샘플은 주로 응답 품질과 수렴 안정성을 결정한다는 것을 관찰했습니다. 이러한 관찰을 바탕으로, 본 논문에서는 엔트로피 기반 적응형 정책 최적화 방법인 EAPO를 제안합니다. EAPO는 현재 정책의 엔트로피와 초기 엔트로피의 비율에 따라 양성 샘플의 가중치 계수를 동적으로 계산합니다. 엔트로피가 감소하는 단계에서는 탐색을 유지하기 위해 양성 샘플에 할당된 가중치를 줄이고, 엔트로피가 증가하는 단계에서는 안정성을 강화하기 위해 가중치를 증폭시켜 엔트로피 붕괴를 완화합니다. 두 개의 공개적으로 사용 가능한 개방형 의료 QA 데이터셋에 대한 실험 결과, EAPO는 응답 다양성과 안정성 측면에서 기존의 고정된 가중치를 사용하는 기준 모델보다 일관되고 크게 우수한 성능을 보였습니다.

Original Abstract

Large Reasoning Models are typically trained via reinforcement learning from verifiable rewards (RLVR). However, existing approaches adopt fixed weights for positive and negative samples, and the conclusions hardly generalize to open-ended question answering (QA). In this paper, we systematically investigate the roles of positive and negative samples in reinforcement learning for open-ended QA. We propose a reward-mean-based strategy for distinguishing positive from negative samples, and observe that negative samples predominantly govern response diversity and the performance upper bound, whereas positive samples primarily determine response quality and convergence stability. Building on these observations, we propose EAPO, an Entropy-driven Adaptive Policy Optimization method that adaptively computes the weighting coefficients of positive samples based on the ratio of the current policy entropy to the initial entropy. During the entropy-decreasing phase, the weight assigned to positive samples is reduced to preserve exploration, whereas during the entropy-increasing phase it is amplified to reinforce stability, thereby mitigating entropy collapse. Experiments on two publicly available open-ended medical QA datasets demonstrate that EAPO consistently and substantially outperforms fixed-weight baselines in both response diversity and stability.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!