MedCalc-R1: 지식 기반 보상 체계 - 의료 수학적 추론을 위한 프레임워크
MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning
수학적 추론 작업을 위한 강화 학습 (RL)에서 검증 가능한 보상(Verifiable Rewards, RLVR) 프레임워크는 일반적으로 허용 오차를 기반으로 실수의 결과를 평가합니다. 그러나 이러한 전략은 임계값 설정의 어려움, 불안정한 훈련 과정, 그리고 특히 임상 시나리오에서 제한된 정확성과 같은 문제점을 가지고 있습니다. 이러한 한계를 해결하기 위해, 우리는 지식 기반 하이브리드 보상 프레임워크인 extsc{MedCalc-R1}을 제안합니다. 구체적으로, 우리는 계산 공식을 명시적으로 생성하도록 강제하는 지식 검증 보상 메커니즘을 도입하여 해석 가능성과 추론의 신뢰성을 향상시키고, 외부 검증기를 사용하여 이를 검증합니다. 또한, 임상 안전 기준에 기반한 하드 제약과 허용 범위 내에서 점진적인 학습을 유도하는 정밀도 민감형 소프트 보상을 결합한 하이브리드 소프트-하드 보상 체계를 설계했습니다. 실험 결과는 우리의 방법이 기존의 기본 모델보다 추론 정확성과 일반화 능력 모두에서 크게 우수하며, 안전이 중요한 영역에서의 효과와 적용 가능성을 입증합니다.
In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated using a tolerance-based reward. However, this strategy suffers from challenges such as difficulty in threshold calibration, unstable training dynamics, and limited accuracy, especially in clinical scenarios. To address these limitations, we propose a knowledge-guided hybrid reward framework (\textsc{MedCalc-R1}). Specifically, we introduce a knowledge verification reward mechanism that enforces explicit generation of computational formulas, which are further validated by an external verifier to enhance interpretability and reasoning reliability. Furthermore, we design a hybrid soft-hard reward scheme combining a hard constraint based on clinical safety thresholds with a soft, precision-sensitive reward that progressively guides learning within the acceptable range. Experimental results demonstrate that our method significantly outperforms existing baselines in both reasoning accuracy and generalization capability, validating the effectiveness and applicability in safety-critical domains.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.