2607.08625v1 Jul 09, 2026 cs.AI

환자 중심 대화형 인공지능의 복잡성

The complexities of patient-centred conversational artificial intelligence

Jonathan Amar
Jonathan Amar
Citations: 26
h-index: 4
João Matos
João Matos
Citations: 21
h-index: 2
Olivia Buege
Olivia Buege
Citations: 0
h-index: 0
Donny Cheung
Donny Cheung
Citations: 0
h-index: 0
Gary S. Collins
Gary S. Collins
Citations: 129
h-index: 5
P. Dhiman
P. Dhiman
Citations: 5,444
h-index: 25
Na-Na Li
Na-Na Li
Citations: 0
h-index: 0
Bingyu Mao
Bingyu Mao
Citations: 67
h-index: 3
Benjamin W Nelson
Benjamin W Nelson
Citations: 38
h-index: 3
Michail Ouroutzoglou
Michail Ouroutzoglou
Citations: 39
h-index: 3
Paul Varghese
Paul Varghese
Citations: 161
h-index: 5

최근 대규모 언어 모델(LLM)을 기반으로 하는 건강 관련 챗봇이 증상 평가에 점점 더 많이 사용되고 있습니다. 그러나 챗봇 개발 및 평가는 종종 협조적이고 명확하게 표현하는 시뮬레이션된 환자를 기반으로 합니다. 본 연구에서는 실제 환자와 챗봇 간의 2,053건의 대화 데이터를 분석한 결과, 사용자마다 의사소통 패턴과 감정 표현이 매우 다양하다는 것을 확인했습니다. 우리는 임상 내용, 정서 상태, 대화 전략 및 의사소통 스타일을 개별적으로 모델링하는 환자 시뮬레이터를 개발했습니다. 15명의 평가자가 참여한 튜링 테스트 방식의 현실성 평가에서, 시뮬레이션된 대화는 실제 대화와 거의 구별할 수 없었으며, 평가자의 정확도는 55%였습니다. 또한, 1,164건의 임상의사가 평가한 사례를 바탕으로 다섯 가지 다양한 환자 유형을 사용하여 네 가지 LLM의 응급 상황 판단 성능을 평가했습니다. 분석 결과, 의사소통 스타일이 삼담 결과를 크게 변화시킬 수 있음을 확인했습니다. 환자 중심 대화형 인공지능은 의사소통 방식의 다양성을 고려해야 합니다. 이상적인 상호작용에 맞춰 설계된 시스템은 실제 환경에서 적용될 경우 성능 저하를 초래하고 건강 불평등을 심화시킬 위험이 있습니다.

Original Abstract

Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%. We used five distinct patient personae, across 1,164 clinician-graded cases, to evaluate the performance of four LLMs in urgency assessment. We found that communication style can significantly alter triage outcomes. Patient-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world.

0 Citations
0 Influential
12.5 Altmetric
62.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!