임상적 상황에서의 대화를 통한 측정: 관찰 가능한 단계 구조, 부분적으로 관찰 가능한 환자 상태
Conversation as Measurement in Clinical Encounters: Observable Phase Structure, Partially Observable Patient State
많은 현대 AI 시스템은 인간의 상호 작용 및 상태를 추론하기 위해 대화 기록을 분석하며, 이러한 정보는 대화에서 추출될 수 있다는 전제를 암묵적으로 가지고 있습니다. 본 연구에서는 '관찰 가능성'에 대해 탐구합니다. 즉, 대화 기록만으로 특정 정보를 얼마나 정확하게 파악할 수 있는지를 다룹니다. 관찰 가능성을 평가하는 것은 어렵습니다. 왜냐하면 기록은 많은 대상에 대한 부분적인 정보만을 제공할 수 있으며, 대규모 분석에는 모델 기반 주석이 필요하기 때문에, 대화 신호의 실제 한계와 주석자의 오류를 구별하기 어렵기 때문입니다. 따라서 본 연구에서는 환자가 직접 보고하는 결과 지표(PROMs)가 환자 상태에 대한 외부적인 기준을 제공하고, 방문이 일관된 구조를 따르는 임상적 상황을 분석합니다. 439건의 실제 임상 기록과 134시간 분량의 대화 내용을 분석하여, 환자 상태와 대화 단계 구조의 관찰 가능성을 연구했습니다. 여기에는 이비인후과 관련 245건의 기록과 이를 보완하는 273개의 PROM 설문이 포함되었습니다. 환자 상태는 발성, 기침, 삼키기와 관련된 PROM 점수를 사용하여 정의하고, 대화 단계 구조는 대화 단계 분할을 통해 분석했습니다. 이러한 분석 결과를 신뢰성 있게 만들기 위해, 개인 식별 정보 보호 규정을 준수하는 GPT-5 모델을 사용하여 기록에 주석을 달고, 40시간 동안 수동 검증을 수행하여, 관찰 가능성의 한계가 단순히 주석자의 오류에서 비롯된 것일 가능성을 줄였습니다. 주요 결과는 '관찰 가능성 불균형'입니다. 즉, 대화 단계 구조는 관찰 가능하며 임상적 상황의 조직을 특징짓는 데 유용하지만, 환자 상태는 부분적으로만 관찰 가능합니다. 이는 환자의 증상과 경험을 이끌어내기 위해 설계된 환경에서도 마찬가지이며, 따라서 대화 기록만을 사용하여 인간의 상태를 추론하는 데 주의가 필요하다는 점을 시사합니다.
Many modern AI systems analyze conversational traces to infer aspects of human interaction and state, implicitly assuming that such information is recoverable from conversation. We study observability: whether a target is recoverable from conversational transcripts alone. Observability is difficult to assess because transcripts may provide only a partial view of many targets, and large-scale analysis requires model-based annotation, making true limits of the conversational signal hard to distinguish from annotator error. We therefore study clinical encounters, where patient-reported outcome measures (PROMs) provide an external anchor for patient state, and visits follow broadly structured patterns. We study observability of patient state and conversational phase structure using 439 real-world clinical encounter transcripts spanning 134 hours, including 245 ENT transcripts paired with 273 PROM surveys. We operationalize patient state using PROM scores for voice, cough, and swallowing; phase structure using conversational phase segmentation. To make these analyses credible at scale, we use a PHI-compliant GPT-5 deployment for transcript annotation and conduct 40 hours of manual validation, reducing the risk that apparent limits of observability simply reflect annotator error. Our core finding is an observability asymmetry: phase structure is observable and useful for characterizing clinical encounter organization, while patient state is only partially observable, even in a setting designed to elicit patient symptoms and experiences, cautioning against transcript-only inference of human state.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.