MindGuard: 멀티턴 정신 건강 지원을 위한 가드레일 분류기
MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support
대규모 언어 모델이 정신 건강 지원에 점점 더 많이 사용되고 있지만, 대화의 일관성만으로는 임상적 적절성을 보장할 수 없습니다. 기존의 범용 안전 장치는 종종 치료적 토로와 실제 임상적 위기를 구별하지 못해 안전 실패를 초래합니다. 이러한 문제를 해결하기 위해, 본 논문에서는 박사급 심리학자들과 협력하여 개발한 임상 기반 위험 분류 체계를 소개합니다. 이 체계는 안전하고 위급하지 않은 치료적 대화는 허용하면서 조치가 필요한 위험(예: 자해 및 타해)을 식별합니다. 또한 임상 전문가가 턴(turn) 단위로 주석을 단 실제 멀티턴 대화 데이터셋인 MindGuard-testset을 공개합니다. 제어된 두 에이전트 설정을 통해 생성된 합성 대화를 사용하여, 경량 안전 분류기 제품군(4B 및 8B 파라미터)인 MindGuard를 학습시켰습니다. 본 연구의 분류기는 높은 재현율(recall) 구간에서 긍정 오류(false positive)를 줄이며, 임상 언어 모델과 결합했을 때 범용 안전 장치에 비해 적대적 멀티턴 상호작용에서 더 낮은 공격 성공률과 유해 관여율을 달성합니다. 우리는 모든 모델과 인간 평가 데이터를 공개합니다.
Large language models are increasingly used for mental health support, yet their conversational coherence alone does not ensure clinical appropriateness. Existing general-purpose safeguards often fail to distinguish between therapeutic disclosures and genuine clinical crises, leading to safety failures. To address this gap, we introduce a clinically grounded risk taxonomy, developed in collaboration with PhD-level psychologists, that identifies actionable harm (e.g., self-harm and harm to others) while preserving space for safe, non-crisis therapeutic content. We release MindGuard-testset, a dataset of real-world multi-turn conversations annotated at the turn level by clinical experts. Using synthetic dialogues generated via a controlled two-agent setup, we train MindGuard, a family of lightweight safety classifiers (with 4B and 8B parameters). Our classifiers reduce false positives at high-recall operating points and, when paired with clinician language models, help achieve lower attack success and harmful engagement rates in adversarial multi-turn interactions compared to general-purpose safeguards. We release all models and human evaluation data.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.