Mental-R1: LLM 추론을 정신 건강 평가에 맞게 조정하기
Mental-R1: Aligning LLM Reasoning for Mental Health Assessment
불안, 우울증 및 자살과 같은 정신 건강 문제는 시급한 전 세계적인 과제이며, 효과적인 개입을 위해서는 적시에 정확한 평가가 중요합니다. 최근에는 대규모 언어 모델이 정신 건강 평가에 활용되고 있습니다. 그러나 기존의 일반적인 사후 훈련 방법은 인간의 평가 과정과 일치하지 않아 신뢰할 수 없는 추론 결과를 초래할 수 있습니다. 이러한 격차를 해소하기 위해, 본 연구에서는 정신 건강 분야에 특화된 강화 학습 프레임워크인 인지 상대 정책 최적화(Cognitive Relative Policy Optimization, CRPO)를 제안합니다. CRPO는 그룹 상대 정책 최적화를 확장하여 정책 최적화 과정에 단계별 불확실성 모델링을 통합합니다. 특히, 우리는 단계별 엔트로피 정규화 메커니즘을 도입하여 초기 추론 단계에서는 광범위한 탐색을 장려하고 후반 단계에서는 확신 있는 의사 결정을 점진적으로 강화함으로써 인간의 인지적 변화를 모방합니다. 또한, 인지 평가 이론에서 영감을 받아 인지적 추론 단계를 공식화하여 이론에 기반한 해석 가능한 추론을 유도합니다. 8개의 정신 건강 데이터 세트에 대한 실험 결과, CRPO는 최고 성능의 강화 학습 기준 모델 대비 가중 F1 점수 측면에서 평균 10.4% 포인트 향상을 달성했습니다. 또한, CRPO로 훈련된 모델 Mental-R1은 기존의 대규모 언어 모델에 비해 추론이 필요한 경우에 명확한 장점을 보여주며, 이는 CRPO가 정신 건강 평가를 위한 추론 능력을 향상시킨다는 것을 시사합니다.
Mental health problems such as anxiety, depression, and suicide remain urgent global challenges, where timely and accurate assessment is critical for effective intervention. Recently, large language models have been explored for mental health assessment. However, existing general-purpose post-training methods do not align with the cognitive processes of human assessment, which may lead to unreliable reasoning outcomes. To bridge this gap, we propose Cognitive Relative Policy Optimization (CRPO), a reinforcement learning framework tailored for the mental health domain. CRPO extends group relative policy optimization by integrating stage-dependent uncertainty modeling into the policy optimization process. Specifically, we introduce a stage-wise entropy regularization mechanism that encourages broad exploration in early reasoning phases and progressively enforces confident decision-making in later stages, mimicking the human cognitive shift from uncertainty to certainty. In addition, inspired by cognitive appraisal theory, we formalize cognitive reasoning stages, thereby guiding theory-grounded interpretable inference. Experiments on 8 mental health datasets show that CRPO achieves an average improvement of 10.4 percentage points in weighted F1-score over the best reinforcement learning baseline. Furthermore, the CRPO-trained model Mental-R1 demonstrates clear advantages compared with existing large language models on reasoning-intensive cases, suggesting that CRPO enhances reasoning capabilities for mental health assessment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.