2605.06191v1 May 07, 2026 cs.AI

퇴원 기록에서 임상적 조치 추출을 위한 대규모 언어 모델의 체계적인 평가

Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction

Ananya Mantravadi
Ananya Mantravadi
Citations: 59
h-index: 3
Shivali Dalmia
Shivali Dalmia
Citations: 6
h-index: 2
P. Desikan
P. Desikan
Citations: 553
h-index: 11

본 논문에서는 CLIP 퇴원 기록 데이터셋을 사용하여 안전에 중요한 임상적 조치 추출을 위한 제로샷 및 퓨샷 대규모 언어 모델(LLM)을 평가합니다. 특히, 환자 이송 과정 및 퇴원 후 환자 안전에 중점을 둡니다. 임상 기록의 복잡성을 관리하기 위해, 서술형으로 작성된 퇴원 기록을 세 단계로 나누어, 정교하고 명확하게 실행 가능한 임상적 작업으로 분해하는 2단계 추출 프레임워크를 제안합니다. 본 논문의 기여는 다음과 같습니다. 첫째, 임상적 조치 추출을 위한 생성형 LLM에 대한 체계적인 평가; 둘째, 범용 LLM과 특정 작업에 특화된 지도 학습 기반 BERT 모델 간의 상세한 비교; 셋째, 다양한 조치 범주 간의 주석 불일치 분석입니다. 실험 결과, 최신 LLM은 지도 학습 모델과 유사하거나 더 나은 성능을 보이는 이진 조치 가능성 감지 능력을 보여주었습니다. 반면, 지도 학습 기반 모델은 특정 작업에 대한 추가 학습 없이, 엄격한 데이터 개인 정보 보호 제약 조건 하에서도, 정교한 다중 레이블 범주 분류에서 여전히 의미 있는 장점을 유지합니다. 질적 오류 분석 결과, 많은 실패는 모델 추론과 데이터셋 주석 규칙 간의 불일치에서 비롯되며, 특히 암묵적인 임상적 조치 및 엄격한 구조적 레이블링 규칙이 적용된 경우에 해당합니다. 이러한 결과는 보고된 성능이 모델의 임상적 추론 능력 부족으로 인해 제한될 수 있음을 시사하며, 이는 단순한 주석으로는 파악하기 어렵습니다. 이유 설명이 없는 레이블은 모델의 임상적 추론 실패와 주석 규칙 불일치를 구별하는 것을 불가능하게 만듭니다. 임상 자연어 처리 기술의 발전에는 특정 구간이 실행 가능하다는 사실뿐만 아니라, 왜 특정 구간이 실행 가능한지에 대한 이유를 기록하는 추론 기반의 데이터셋이 필요하며, 이를 통해 모델의 임상적 이해도를 적절하게 평가할 수 있습니다.

Original Abstract

The work in this paper evaluates zero-shot and few-shot large language models (LLMs) for safety-critical clinical action extraction using the CLIP discharge-note dataset, with particular emphasis on transitions of care and post-discharge patient safety. To manage the complexity of clinical documentation, we introduce a two-stage extraction framework that decomposes discharge notes, that are written in narrative form, into fine-grained, explicitly actionable clinical tasks through a staged prompting strategy. Our contributions include a systematic assessment of generative LLMs for clinical action extraction, a detailed comparison between general-purpose LLMs and task-specific supervised BERT-based models, and an analysis of annotation inconsistencies across different action categories. We show that contemporary LLMs achieve performance comparable to or exceeding supervised models on binary actionability detection, while supervised baselines retain a meaningful advantage on fine-grained multi-label category classification, despite the absence of task-specific fine-tuning and under strict data-privacy constraints. Qualitative error analysis reveals that many failures stem from misalignment between model reasoning and dataset annotation conventions, particularly in cases involving implicit clinical actions and rigid structural labeling rules. These results indicate that reported performance reflects model limitations due to lack of clinical reasoning, that is not captured by plain annotations. Labels without rationales make it impossible to distinguish clinical reasoning failures from annotation convention mismatches. Advancing clinical NLP requires reasoning-annotated datasets that document why specific spans are actionable, not merely which spans were labeled, enabling proper evaluation of model clinical understanding.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!