SIMAX: 다중 신뢰도 및 주석이 포함된 의료 전문가-환자 대화 시뮬레이션을 위한 확장 가능하고 해석 가능한 프레임워크
SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation
배경: 앰비언트 디지털 기록 시스템의 광범위한 도입으로 인해 의료 전문가와 환자의 대화 데이터가 대량으로 수집되고 있습니다. 임상 커뮤니케이션 데이터에 대한 인간 코딩은 비용이 많이 들고 일관성이 부족하며 확장하기 어렵기 때문에 AI 기반 커뮤니케이션 코딩 시스템 개발의 필요성이 제기되었습니다. 그러나 이러한 시스템을 평가하려면 실제 대화와 인간이 코딩한 레이블이 필요하지만, 이는 대규모로 얻기가 어렵습니다. 방법: 우리는 SIMAX(Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation)라는 프레임워크를 개발했습니다. 이 프레임워크는 참조 행동 주석과 함께 제어된 임상 대화 데이터를 생성합니다. SIMAX는 미리 정의된 임상 시나리오, 페르소나 및 음성 조건을 기반으로 의료 전문가와 환자의 대화를 생성하며, 목표 커뮤니케이션 행동을 설정합니다. 행동은 두 가지 코드북을 사용하여 제어됩니다. 하나는 전체적인 커뮤니케이션 품질을 위한 글로벌 코드북이고, 다른 하나는 특정 측정 가능한 행동을 위한 WISER 코드북입니다. 우리는 자동 및 인간 평가를 통해 SIMAX를 평가했으며, 또한 예시 커뮤니케이션 코딩 시스템에 대한 평가도 수행했습니다. 결과: SIMAX는 세 가지 전문 분야, 다양한 방문 단계, 페르소나 특성 및 억양 조건에서 총 3,388개의 시뮬레이션 대화를 생성했습니다. 자동 평가는 평균 UTMOS 및 WV-MOS 점수가 각각 3.03과 2.61이며, WER 및 CER는 각각 0.07과 0.05이고, CLAP 코사인 유사도는 0.41임을 보여주었습니다. 이는 음성의 자연스러움이 합리적이고, 전사 정확도가 높으며, 텍스트와 오디오가 잘 일치한다는 것을 나타냅니다. 인간 평가는 중간 MOS 점수가 4.67이고, 임상 현실성 점수는 3.00이었습니다. 추가적인 평가 결과, SIMAX는 커뮤니케이션 코딩 시스템이 행동 목표에 어떻게 반응하는지 평가하고 특정 측면에서 민감도가 부족함을 보여줄 수 있습니다. 결론: SIMAX는 제어되고 재현 가능한 시뮬레이션 의료 전문가-환자 대화를 생성하여, 커뮤니케이션 코딩 시스템을 개발, 검증 및 개선하기 위한 데이터 기반을 제공합니다.
Background. The widespread deployment of ambient digital scribes is driving large-scale capture of clinician-patient dialogues. Human coding of clinical communication data remains costly, inconsistent, and difficult to scale, motivating AI-driven communication coding systems. However, evaluating these systems requires real-world dialogues and human-coded labels, both hard to obtain at scale. Methods. We developed SIMAX (Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation), a framework for generating controlled clinical dialogue data with reference behavioral annotations. SIMAX generates clinician-patient dialogues from predefined clinical scenarios, personas and voice conditions, and target communication behaviors. Behaviors are controlled using two codebooks: the Global Codebook for overall communication quality and the WISER Codebook for specific countable behaviors. We evaluated SIMAX using automated and human quality assessments and an example communication coding system. Results. SIMAX generated 3,388 simulated dialogues across three specialties, multiple visit stages, persona characteristics, and accent conditions. Automated assessment showed mean UTMOS and WV-MOS scores of 3.03 and 2.61, WER and CER of 0.07 and 0.05, and CLAP cosine similarity of 0.41, suggesting reasonable speech naturalness, high transcription fidelity, and positive text-audio correspondence. Human evaluation showed a median MOS of 4.67 and a median clinical realism score of 3.00. Downstream evaluation suggests that SIMAX can assess how a communication coding system responds to behavioral targets and reveal insufficient sensitivity in some dimensions. Conclusions. SIMAX generates controlled and reproducible simulated clinician-patient dialogues, providing a data foundation for developing, validating, and refining communication coding systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.