2602.00950v1 Feb 01, 2026 cs.AI

MindGuard: 멀티턴 정신 건강 지원을 위한 가드레일 분류기

MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support

António Farinhas
António Farinhas
Citations: 72
h-index: 5
Nuno M. Guerreiro
Nuno M. Guerreiro
Instituto de Telecomunicações
Citations: 1,990
h-index: 19
José Pombal
José Pombal
Sword Health
Citations: 751
h-index: 10
P. Martins
P. Martins
Citations: 490
h-index: 6
Alex Conway
Alex Conway
Citations: 3
h-index: 1
Cara Dochat
Cara Dochat
Citations: 2
h-index: 1
Maya D'Eon
Maya D'Eon
Citations: 7
h-index: 2
Ricardo Rei
Ricardo Rei
Citations: 430
h-index: 11
L. Melton
L. Melton
Citations: 2
h-index: 1

대규모 언어 모델이 정신 건강 지원에 점점 더 많이 사용되고 있지만, 대화의 일관성만으로는 임상적 적절성을 보장할 수 없습니다. 기존의 범용 안전 장치는 종종 치료적 토로와 실제 임상적 위기를 구별하지 못해 안전 실패를 초래합니다. 이러한 문제를 해결하기 위해, 본 논문에서는 박사급 심리학자들과 협력하여 개발한 임상 기반 위험 분류 체계를 소개합니다. 이 체계는 안전하고 위급하지 않은 치료적 대화는 허용하면서 조치가 필요한 위험(예: 자해 및 타해)을 식별합니다. 또한 임상 전문가가 턴(turn) 단위로 주석을 단 실제 멀티턴 대화 데이터셋인 MindGuard-testset을 공개합니다. 제어된 두 에이전트 설정을 통해 생성된 합성 대화를 사용하여, 경량 안전 분류기 제품군(4B 및 8B 파라미터)인 MindGuard를 학습시켰습니다. 본 연구의 분류기는 높은 재현율(recall) 구간에서 긍정 오류(false positive)를 줄이며, 임상 언어 모델과 결합했을 때 범용 안전 장치에 비해 적대적 멀티턴 상호작용에서 더 낮은 공격 성공률과 유해 관여율을 달성합니다. 우리는 모든 모델과 인간 평가 데이터를 공개합니다.

Original Abstract

Large language models are increasingly used for mental health support, yet their conversational coherence alone does not ensure clinical appropriateness. Existing general-purpose safeguards often fail to distinguish between therapeutic disclosures and genuine clinical crises, leading to safety failures. To address this gap, we introduce a clinically grounded risk taxonomy, developed in collaboration with PhD-level psychologists, that identifies actionable harm (e.g., self-harm and harm to others) while preserving space for safe, non-crisis therapeutic content. We release MindGuard-testset, a dataset of real-world multi-turn conversations annotated at the turn level by clinical experts. Using synthetic dialogues generated via a controlled two-agent setup, we train MindGuard, a family of lightweight safety classifiers (with 4B and 8B parameters). Our classifiers reduce false positives at high-recall operating points and, when paired with clinician language models, help achieve lower attack success and harmful engagement rates in adversarial multi-turn interactions compared to general-purpose safeguards. We release all models and human evaluation data.

2 Citations
0 Influential
9.5 Altmetric
49.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!