2607.08257v1 Jul 09, 2026 cs.AI

MentalHospital: 정신과 임상 상담 평가를 위한 가상 환경

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

Yuming Yang
Yuming Yang
Citations: 13
h-index: 2
Xiao Sun
Xiao Sun
Citations: 58
h-index: 2
Jiang Zhong
Jiang Zhong
Citations: 32
h-index: 3
Haoyang Zeng
Haoyang Zeng
Citations: 4
h-index: 1
Jingwang Huang
Jingwang Huang
Citations: 16
h-index: 2
Kaiwen Wei
Kaiwen Wei
Citations: 6
h-index: 2
Yuanwei Zou
Yuanwei Zou
Citations: 0
h-index: 0
Yun Chen
Yun Chen
Citations: 0
h-index: 0
Zhengxiao Wu
Zhengxiao Wu
Citations: 0
h-index: 0

대규모 언어 모델(LLM)은 대화, 진단 및 치료 계획을 포함한 다양한 정신과 관련 작업에서 뛰어난 성능을 보여주었지만, 기존의 벤치마크는 실제 정신과 임상 상담 과정을 충분히 반영하지 못하는 경우가 많습니다. 본 연구에서는 LLM 기반의 정신과 임상 상담 평가를 위한 가상 환경인 $ extbf{MentalHospital}$을 소개합니다. MentalHospital은 주관적 면담, 객관적 검사, 진단 평가 및 치료 계획 (S.O.A.P.) 워크플로우를 구현하며, 1,193건의 익명화된 정신과 전자 건강 기록(EHR) 데이터로부터 구축된 표준 환자를 활용하여 ICD-11의 주요 범주와 76가지 질환을 포괄합니다. 각 상담은 객관적인 EHR 기반 기준과의 비교 분석과 함께 임상 과정의 품질에 대한 주관적 평가를 결합한 이중 추적 프로토콜을 통해 평가됩니다. 전문적인 판단의 신뢰도를 높이기 위해, $ extbf{MentalEval}$이라는 5가지 영역별 평가 도구를 개발했습니다. MentalEval은 의사소통 공감 능력, 면담 전문성, 임상 기록 품질, 진단 정확성 및 치료 적절성에 대한 평가를 수행하며, 명확한 평가 기준과 전문가 지도를 통해 학습된 SFT(Supervised Fine-Tuning) 및 DPO(Direct Preference Optimization) 방식으로 훈련되었습니다. 22명의 임상의로부터 수집된 설문조사 결과는 MentalHospital의 실제 임상과의 유사성(3.88/5)을 나타내며, MentalEval은 평균 QWK 점수 0.944로 전문가 간 높은 일치도를 보입니다. 벤치마크 테스트 결과, 현재까지 개발된 가장 강력한 LLM도 객관적인 정신과 능력 측면에서 임상의보다 37.28%p 뒤처지는 것으로 나타났으며, 특히 정신 상태 평가가 주요 성능 저하 요인으로 확인되었습니다.

Original Abstract

Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, and Treatment Planning (S.O.A.P.) workflow, using skill-augmented standardized patients constructed from 1,193 de-identified psychiatric electronic health record (EHR) cases spanning all major ICD-11 categories and 76 disorders. Each encounter is assessed through a dual-track protocol that combines objective comparison against EHR-derived references with subjective assessment of clinical process quality. To scale specialist judgment, we develop $\textbf{MentalEval}$, five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO. Survey responses from 22 clinicians support MentalHospital's clinical fidelity (3.88/5), while MentalEval achieves strong expert alignment with an average QWK of 0.944. Benchmarking shows that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment as a key bottleneck.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!